AI Agents / POV
System One Models: Jev vs Laya and How They Change AI Agents
System One models return fast, structured decisions instead of prose. Learn how Jev and Laya compare and where they fit in AI agents.
On this page
- Why call them “System One” models?
- How small decisions become an agent
- What do Jev and Laya return?
- Jev vs Laya: what changed when I trained Laya?
- Is this just classification?
- RLCD, RLHF, and uncertainty
- Is Jev actually cheap?
- Where should code or a generative model take over?
- What this changes for agents
- Sources
Most AI agents are built around a generative model. We ask it to understand a request, choose a tool, judge the tool’s result, decide whether to continue, and write an answer. Generative models are powerful, but many of those steps require a small decision rather than a written response.
Consider a customer message: “I have been charged twice. Please fix this today.” Before drafting a reply, a support system needs to decide whether this is a billing issue, how urgent it is, and whether a refund request needs review. The application may need a category, a score, or a probability. It does not necessarily need a paragraph explaining each judgment.
That is the appeal of System One models: they take context and bounded questions, then return structured decisions that software can use. Jev is TypeSafe AI’s hosted model. Laya is an open-weight implementation of a similar approach.
Why call them “System One” models?
The name draws on Daniel Kahneman’s distinction between fast, intuitive System 1 thinking and slower, deliberate System 2 thinking. The analogy has limits: AI models are not human minds. But it suggests a useful division of work.
“Billing, urgent” is a quick judgment about the customer’s message. Checking whether there was actually a duplicate charge, applying a refund policy, and explaining the outcome require evidence and more deliberate reasoning.
An agent makes many such small judgments: Which tool should it use? Is a retrieved passage relevant? Should it search again? Does a proposed answer need human review? A decision model can handle a well-scoped judgment while code manages exact rules and a generative model handles reasoning and language.
How small decisions become an agent
Computers build complex behaviour from simple operations. A program repeatedly checks conditions (1 and 0), takes branches, and combines the results. The challenge is that real-world conditions are often ambiguous. We can write:
if customer_is_likely_to_cancel:
alert_retention_team
But how do we calculate customer_is_likely_to_cancel when the customer never says “cancel”? A keyword rule is too rigid. A generative model can judge it, but asking one to produce and parse a written response at every branch may be unnecessary. A decision model can estimate the answer; code can decide what to do with that estimate.
This is the architecture I find interesting: many focused judgments connected by explicit logic. A generative model remains part of the system, but it is called when its broader abilities are needed.
Imagine an agent answering a question about company policy:
- Code retrieves candidate documents.
- A decision model scores which passages are relevant.
- A generative model reasons across the selected passages and writes an answer.
- A focused check flags claims that the cited passages may not support.
- Uncertain or consequential answers go to a person.
This design makes each judgment visible and testable. It also creates dependencies: if the relevance check discards the only useful passage, the writing model cannot recover that evidence. Any production workflow needs end-to-end evaluation.
What do Jev and Laya return?
TypeSafe defines three question types. Its documentation says they can be asked together against the same input state.
| Question type | Example | Result |
|---|---|---|
| Choice | Which team should handle this ticket? | One option from a defined list, plus probabilities for the options |
| Score | How urgent is the ticket? | A rating against defined levels, plus their probabilities |
| Noul | Does the message request a refund? | A probability that the answer is yes |
Jev accepts text and text-bearing structured input such as a JSON object. It does not directly accept images, audio, or video, according to its model reference. Laya provides a compatible style of typed decision-making with downloadable weights and a fine-tuning workflow.
A constrained output is dependable in one specific sense: a Choice is selected from the options supplied. Valid structure does not guarantee a correct judgment. If the options omit the right answer, or the model misunderstands indirect wording, it can still return a confident mistake. Applications should include an “other” option when appropriate and test decisions on representative cases.
Jev vs Laya: what changed when I trained Laya?
In my own use case, Jev initially made more accurate decisions than the Laya checkpoint I started with. After I trained Laya for those decisions, Laya performed better than Jev in my tests.
That is an observation from a task-specific comparison, not a general benchmark. Without a published test set, sample size, metric, and evaluation procedure, it should not be read as evidence that one model is broadly more accurate.
| Dimension | Jev | Laya |
|---|---|---|
| Access | Hosted API | Open weights; can run locally or be hosted |
| Adaptation | Change the state, questions, instructions, and criteria in each request | Change the request and fine-tune model weights |
| Operations | TypeSafe runs the service | You can run and maintain the deployment |
| My task-specific tests | Better before Laya training | Better after training for my use case |
TypeSafe says Jev is not fine-tuned on customer data; users shape its behavior through each request. Laya’s repository publishes weights and fine-tuning instructions. The practical question is which approach performs better on your repeated decisions, under your latency and operating constraints.
Is this just classification?
There is overlap. A classical classifier can return a label and probabilities. With a stable task and good labeled data, logistic regression, a random forest, or a task-specific transformer may be simpler, faster, or more accurate.
The distinction is mainly the interface and intended use. Jev lets an application provide state, define bounded questions and answer options at request time, and receive typed results for several judgments. A dedicated classifier is usually built around a more fixed label set and task.
Neither approach has a monopoly on trustworthy probabilities. Classical classifiers can be calibrated, and a decision model’s probabilities can be poorly calibrated on a particular workload. Test that property with held-out examples from the setting where the system will run.
RLCD, RLHF, and uncertainty
Reinforcement Learning from Human Feedback, or RLHF, is commonly used to steer generated responses toward human preferences. TypeSafe describes Reinforcement Learning for Calibrated Decisions, or RLCD, as training for decisions whose probabilities track observed outcomes. TypeSafe’s introduction presents this as a central part of its approach.
Calibration has a precise meaning. Among many comparable cases assigned roughly 70% probability, the predicted outcome should occur roughly 70% of the time. It is a property of groups of predictions, not a promise about any one case. Scikit-learn’s calibration guide explains how to assess it.
One more distinction matters in application code: Jev’s Choice and Score responses include a confidence field derived from the shape of their probability distributions. That field is not simply the probability that the selected answer is correct. A Noul response gives its yes probability without a separate confidence field. See TypeSafe’s question-type reference.
Is Jev actually cheap?
TypeSafe currently lists Jev at $0.042 per million input tokens, with no charge for output tokens, and reports end-to-end responses of roughly 70–500 ms in its launch material. These are published figures, not guarantees for every deployment. See the model reference and launch post.
The unit of comparison matters. A decision model performs one part of a workflow, and we need to perform multiple decision calls within a same workflow. A generative model might classify, reason, and draft in one call. Comparing the price of a single Jev judgment with a generative call that does more work can exaggerate the saving.
For illustration, suppose a business receives 200,000 customer emails a day, and each Jev request contains 1,000 input tokens. One request per email is 200 million input tokens: about $8.40 per day, or $252 over 30 days, at the listed price. Five separate requests of that size per email would be about $42 per day, or $1,260 over 30 days. Asking several questions in one request may avoid repeating the input.
Those figures cover Jev input tokens only. A fair comparison measures the whole completed workflow: later model calls, infrastructure, review time, latency, accuracy, and the cost of mistakes.
Where should code or a generative model take over?
Use ordinary code for exact rules, such as calculating a refund amount or checking a date range. Use a generative model when the task requires substantial reasoning, writing, or coding. Keep a person in the loop when an incorrect automatic action would be costly.
A decision model is most useful when the question is bounded but the evidence is messy: ticket routing, passage relevance, urgency, or whether a case needs review. Even there, choose automation thresholds from results on your own workload. A probability is information for a policy; it is not permission to act.
What this changes for agents
An agent can be a network of explicit decisions rather than one model asked to handle every detail. Code provides exact rules. Focused models make uncertain judgments. Generative models handle broader reasoning and language. People review the cases whose stakes or uncertainty warrant it.
This will not make agents perfectly deterministic. It does make their decisions easier to inspect, measure, improve, and replace. The most useful question System One models raise is not simply “How fast is one decision?” It is: What decisions is this agent making, what evidence does each use, and how do we know the entire workflow works?