Agents spend most of their cycles on bounded decisions: classifying tickets, routing requests, gating tool calls, scoring responses. Until mid-September, large language models handled these tasks alongside text generation, producing hundreds of tokens to answer questions that amount to yes or no. On September 15, a startup called TypeSafe AI shipped a model built exclusively for bounded decisions, generating zero output tokens and returning typed probabilities instead of prose. Within three weeks, six more companies followed. A new layer of the AI stack has arrived with enough competition, open-weight availability, and speed to reshape how agents are built.
System One
On September 15, TypeSafe AI released Jev, the first model in a category that its founder Diogo Almeida calls “System One,” borrowing Daniel Kahneman’s term for fast, intuitive cognition. Almeida, a former OpenAI researcher who helped develop the reinforcement learning from human feedback technique behind ChatGPT, built Jev to accept unstructured input — text, JSON, program state — alongside a set of typed questions, and return calibrated probabilities over predefined answers in a single parallel pass. The model generates zero output tokens. It returns numerical probabilities that software can branch on directly, with no parsing step between inference and action.
The inefficiency that Jev targets is specific and pervasive. Most production AI workloads involve bounded decisions: classifying a support ticket, routing a request to the appropriate model tier, deciding whether to invoke a tool. These decisions have been handled by large language models generating prose that then requires parsing back into structured data. That workflow consumes seconds per call, bills by output token for answers that amount to a single bit of information, and breaks when the model returns malformed JSON or hallucinates an option outside the predefined list.
Jev’s pricing reflects the efficiency gain: $0.042 per million input tokens, with output tokens billed at zero. TypeSafe reports end-to-end latency between 70 and 500 milliseconds, which the company claims represents a 40x to 200x speedup over frontier LLMs on comparable classification tasks. TypeSafe launched with a $40 million seed round led by DCVC and published nine known failure modes on launch day, documenting the model’s boundaries alongside its capabilities.
The bandwagon gets rolling
The category filled out in seventeen days. Fastino Labs shipped GLiNER2.5-Decide on September 24, a compact 340-million-parameter encoder built for local CPU inference. Liquid AI followed on September 29 with d1, a hosted model that claimed the top position on Hugging Face’s Decision Index across 40 benchmarks. OpenAI previewed a Decisions API at DevDay the same day, constraining GPT-6 Luna to developer-defined questions with finite answers. On October 1, Cloudflare, Perplexity, and AWS all shipped decision models within hours of one another.
The architectural approaches diverged sharply. Cloudflare fine-tuned Qwen 3.8-27B with rank-256 LoRA adapters and released the weights under Apache 2.0. Its smaller Clef-flash variant, built on Qwen 3.5-9B, delivers a median decision latency of 38.8 milliseconds. Perplexity took a similar path, open-sourcing pplx-decider-v1-27b at $0.04 per million input tokens with free output. AWS chose a fundamentally different scale with Strands Decider 2B, stripping the text-generation head from Qwen3.5-2B entirely and replacing it with a pointer head of roughly one million parameters. The result runs locally on a single consumer GPU in about 115 milliseconds. Liquid AI and OpenAI kept their weights proprietary.
OpenAI’s approach stood apart from the purpose-built entries. Its Decisions API constrains GPT-6 Luna, the company’s lowest-cost model, to developer-defined questions with finite answers, claiming a 10x speedup over standard Luna inference, from 1.6 seconds per call down to 150 milliseconds. This repurposes an existing model for decision tasks rather than training a new architecture for them, and OpenAI has released the API only in limited preview.
The two-tier inference stack
The speed of convergence reflects a structural mismatch that the agent industry has tolerated for some time. An AI agent deciding whether to invoke a tool typically routes that decision through a frontier LLM. The model generates hundreds of tokens to produce an answer that amounts to yes or no, consuming seconds and thousands of input tokens per decision. Agents that make dozens of these micro-decisions per task accumulate latency and cost at a rate that scales poorly with workflow complexity.
Decision models create a two-tier inference stack. Bounded choices such as classification, routing, gating, and scoring move to a dedicated layer that returns structured answers at millisecond latency, while the LLM handles generative work that requires composed text. AWS’s Strands Decider documentation makes the division explicit: the model gates each tool call inside an agent loop, answering two yes/no questions about argument validity and premature invocation before the tool runs. The decision model serves as a fast, cheap checkpoint that the agent passes through dozens of times per task, reserving the expensive generative model for the responses that actually require language.
Three of the six new entrants released open weights under Apache 2.0, which shifts the competitive dynamic. Routing logic, the layer that decides how every request flows through an agent system, can now be self-hosted, fine-tuned on proprietary data, and run without API calls. Cloudflare has already announced a reinforcement-learning fine-tuning service for Clef, initially assisted by its engineering team and later planned as a self-serve platform, aimed at companies with enough labeled historical data to train a custom decision model. The open-weight entries also commoditize the category before it has had time to consolidate, pressing proprietary vendors to compete on calibration quality and ecosystem integration rather than on access.
Separation of concerns
The decision model went from a single entrant to seven in seventeen days, with open weights available from at least four. The pattern echoes other layers of the software stack: once a class of work can be cleanly separated from a general-purpose tool, specialized alternatives appear fast and commoditize faster. For agent architectures specifically, the category addresses a bottleneck that will grow more acute as workflows lengthen and the number of micro-decisions per task rises. Open weights, competitive pricing, and sub-40-millisecond latency mean that the decision-model infrastructure already exists to run invisibly in the request path. Building agents without a dedicated decision layer will increasingly require justification.


