Say a support ticket contains one line â âlocked out againâ â with a screenshot of an error pasted underneath. A person usually looks at the image first: is this the login page, the payment page, or did the site not even load? Only then do they know what to ask and who should handle it.
Reading about Cloudflareâs Clef, thatâs the kind of enterprise agent I had in mind: it doesnât necessarily write code, but it may spend all day reading support tickets, looking at attachments, and deciding what happens next. Where those steps only need one of a few known answers, does a decision model have a shot at shortening the path, or speeding up the handling end to end? Thatâs the same question the previous post on Jev asked â except this time Clef puts rather more on the table.
This is a read of where things publicly stand as of October 2, 2026, plus the directions I think are worth trying. I havenât run anything yet, and the workflow examples below are all hypothetical.
Where Clef and Clef-flash currently stand
Cloudflare launched Clef and Clef-flash on October 1, both available through a hosted Workers AI endpoint. (launch post)
Clef is a 27B model post-trained from Qwen3.8-27B, keeping the vision encoder. It reads a state and a question schema, and emits a probability for every allowed answer. The weights, the joint schema head, and the inference code are published under Apache 2.0, so you can deploy it yourself â which is not the same as the full training data and process being public. (Clef model card)
Clef-flash is post-trained from Qwen3.5-9B, a smaller sibling thatâs also multimodal and uses the same decision interface. Cloudflare positions it for latency-sensitive cases. Being the smaller model doesnât let you conclude itâs worse at every task â that still depends on the actual problem. (Clef-flash model card)
For a developer, thatâs two routes rather than one: validate the problem against the hosted API first, or download the model and decide how to deploy it yourself. The second route, of course, brings GPUs, service operations, and capacity planning back along with it.
Which part of the work itâs trying to speed up
Narrow the problem down a bit first. That screenshot doesnât need a detailed diagnostic report right now â the workflow only wants to know what category the problem falls into, whether the image carries enough of a clue, and which queue it should go to.
Clefâs approach is to let the Qwen backbone read the input, then have a dedicated head score the options in the schema. Cloudflare describes this decision step as non-autoregressive: thereâs no intermediate text to generate token by token before you extract an answer from it. (architecture notes)
So what itâs primarily trying to improve is the cost of this kind of fixed-answer-space judgement. Reading the context and the image still takes compute; if fetching the data upstream is slow, swapping in a decision model wonât automatically fix the latency of the whole workflow.
The interface carries over Jev / System Oneâs choice, noul, and score: pick a category, return the probability of yes, rate against an ordered rubric. That gives code that has already decomposed its questions a chance to reuse the structure â but once the model changes, the judgement quality and the thresholds still have to be re-validated. âThe request goes throughâ isnât the thing youâre checking.
What the images open up
Back to the hypothetical support ticket: you could send the text and the screenshot together, ask whether the image contains a login error, whether more information is needed, and then pick a routing result from the known categories. That creates an opening to drop a step â describing the whole image as prose first, then classifying the prose. How much latency actually comes off, and whether important detail gets lost, both still need testing.
That said, model capability and the hosted API have to be read separately right now. The published inference code takes images and video as frame arrays; the schema Workers AI currently lists accepts at most four embedded images, rejects remote image URLs, and lists no video field. Seeing that the model supports video doesnât let you assume the hosted endpoint takes the same call. (local input format, Workers AI parameters)
Another direction worth trying is document pre-processing: what kind of document an attachment is, whether an image is legible, whether a specified field appears on screen â then decide which existing path it goes down. Exact monetary calculation still belongs in code, and anything requiring approval still has to go through the existing permission and review machinery.
Here too, âdata not providedâ has to be an explicit state. If the screenshot simply didnât capture the error message, the reasonable next step is to ask for another one. The model wavering between several categories is more likely a hand-off to a different model or a person. Both pause the workflow, but they need different remedies.
The wild growth that might follow
Cloudflare has announced the service direction where its FDE team assists with fine-tuning, with a self-serve platform planned for later. (service details)
Following that line, my own guess is that companies may gradually turn the judgements they make most often into schemas and datasets they can evaluate repeatedly â which support tickets suit which queue, which attachments require a follow-up request. Then model selection or fine-tuning has a reasonably concrete target, instead of wiring a model into the agent first and hoping it finds its own purpose.
Clef-flash also makes me think about tiering: handle the categories whose performance has been verified as stable, and send everything else to a larger model or a person. Though two models may well make similar mistakes, so whether the hand-off actually improves the outcome is its own question. The second model isnât automatic insurance.
For now Iâd pick one small workflow and validate there. Alongside accuracy, Iâd look at p95 latency, the cost of both missed and wrong classifications, and how many cases end up with a human anyway. Self-hosted and hosted API both need testing under their own real deployment conditions â public benchmark numbers arenât a production commitment.
If those conditions hold, Clef could become a genuinely practical component inside an enterprise agent: continuously reading textual and visual state, and supplying the workflow with judgements whose scope is clearly bounded. Whatâs worth watching next is which jobs it reliably saves time on, and whether the time saved buys more exception handling than itâs worth.