Tool of the Week: Decision models, and the end of parsing a chatbot's answer

A lot of automation has a step that looks like this: read a ticket, an email or a file, then decide which of five buckets it goes in. The usual build is to ask a chat model for the label and then parse its reply. Sometimes it says "Category: Billing." Sometimes it writes a paragraph. Your parser breaks at 2 AM.

On October 1 Cloudflare released Clef and Clef-flash, open-weight models built for exactly this step. You send the content and the answers you will accept. You get back probabilities, not prose. Perplexity's pplx-decider and Amazon's Strands Decider 2B are the same idea from two other vendors.

Details for Clef: it comes in 27B and 9B sizes, reads text plus images and video frames, and takes up to 64 questions per request in three forms (yes or no, pick one option, rate on a scale). Weights are Apache 2.0 and it also runs hosted on Workers AI. Cloudflare reports median decision times of about 39 milliseconds for the small one and 209 for the large one. Those are Cloudflare's numbers, not independent tests. I could not confirm Cloudflare's pricing, so I am not quoting one.

Why this matters if you run IT or data work: a probability can be compared to a threshold. Above 0.9, route it automatically. Between 0.6 and 0.9, send it to a person. Below that, hold it. That is an approval gate you can audit, which a free-text answer never gave you.

Limits are real. Cloudflare's own coverage says Jev still beats Clef on knowledge-heavy tests, so this is for sorting and routing, not for questions that need broad knowledge. Test on your own tickets and files before you trust any benchmark.

Quick Hits

  • Claude Haiku 5.5 cut the small-model price by about 75 percent. Released October 7. API pricing is $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, and $0.50 and $2.50 above that. Anthropic says it costs about 75 percent less to run than Haiku 4.5 and points it at high-volume work like summarization, classification and subagent tasks. Those benchmark gains are Anthropic's own, so treat them as claims until you test on your data.

  • Amazon released Strands Decider 2B, a decision model you host yourself. About 2 billion parameters, Apache 2.0, released October 1. It reportedly runs in around 100 milliseconds on a consumer GPU and is built to route requests, pick tools and review agent actions. Trade coverage notes two cautions: adversarial text in an agent's input can sway it, and Amazon has not shown that hosting it yourself is cheaper than a hosted option once hardware and upkeep are counted.

  • Perplexity opened its own decision model at $0.04 per million input tokens. pplx-decider is a 27B open-weight model under Apache 2.0, with a hosted API that does not charge for output. Same pattern as Clef: fixed answers in, probabilities out. Three vendors shipping this within a week suggests the classify-and-route step is about to get cheap.

Prompt of the Week: The Decision Spec

I have a step in a workflow where something (a ticket, email, file or
record) gets sorted or approved. Help me turn it into a testable
decision.

Step description: [what arrives, who decides today, what happens next]

Do this in order:

1. List the allowed answers. Each one must be a fixed label with a
   one-line definition. If two labels overlap, tell me and fix it.
2. For each label, write two example inputs that clearly belong and
   one that is borderline.
3. Propose confidence thresholds: above what score it runs
   automatically, in what range a person reviews it, and below what
   score it is held.
4. Name the worst wrong answer (the mistake that costs the most) and
   say which label it would be mistaken for.
5. Write 10 test cases I can run, including 3 that try to trick it.

Fill in one real step from your week. If you cannot list the allowed answers in step 1, the step is not ready to automate, and that is useful to know before you build anything.

One last thing

Writing this is one side of the work. The other is building it: data pipelines, file and reporting workflows, integrations between systems that were never meant to talk, and the automation that takes the manual steps out.

If a process in your week still depends on someone remembering to run it, book a free 15-minute audit and I will find your top 3 time-wasters. If it is not worth automating, I will tell you that.

Scott