Home
All posts

· AI research

How to (Not) Reinvent the Wheel With Jeff

Why this mattersTurning a model’s output into a clear decision makes it easier to use and check in a manufacturing workflow.

There is a new model in town that does not talk. It does not write essays, explain its reasoning, or apologize for being late. It returns probabilities over things you already defined.

The model is Jev, from TypeSafe AI. The company presents it as a “System One” model for fast, structured decisions. Independent open-source projects such as Jeff explore a similar interface. The idea is appealing: give software the kind of output its decision needs, instead of turning every judgment into a paragraph [1, 3].

Choice, Score, and Noul

Jev accepts typed questions about a state supplied as text, logs, tickets, or a structured summary. Its three primitives are:

PrimitiveQuestionExample
ChoiceWhich predefined option?Continue / re-measure / review
ScoreWhere on an ordered rubric?Quality or risk level 0–3
NoulHow likely is this statement to be true?Probability that a reading is anomalous

Choice and Score return distributions and confidence; Noul returns a value between zero and one. The official documentation says the questions in one request are evaluated independently and in parallel against the same state. More complex decisions can then combine those answers in ordinary code [2].

This resembles a smart function call. A caller defines the decision space and receives values it can branch on. The output type rules out malformed labels, while the probabilities expose uncertainty for the next step.

A shared body, different decision spaces

A conventional autoregressive language model usually projects its hidden state onto a vocabulary-sized output layer. Prose, code, and a yes/no answer all pass through token prediction. A multi-head decision model can instead reuse a representation and map it into several smaller output spaces:

h=fθ(x),pt(y=k∣x)=softmax⁡(Wth+bt)kh=f_\theta(x),\qquad p_t(y=k\mid x)=\operatorname{softmax}(W_t h+b_t)_k

The encoder fθf_\theta represents the input state xx. The parameters Wt,btW_t,b_t define the output head for task or decision tt. A routing head predicts a categorical distribution; an ordinal head describes levels on a rubric; a binary head estimates the probability of a specified event.

y^choice=arg⁡max⁡kpk,s^=∑k=0K−1vkpk,q=σ(w⊤h+b)\hat y_{\rm choice}=\arg\max_k p_k,\qquad \hat s=\sum_{k=0}^{K-1}v_kp_k,\qquad q=\sigma(w^\top h+b)

Here vkv_k is the value assigned to rubric level kk, and σ\sigma is the logistic sigmoid. The expected score retains information that is lost by selecting only the most likely level. A binary probability qq can feed an anomaly or review rule.

These equations describe a concrete implementation of the design idea. Jev’s public materials establish its typed interface and parallel evaluation, but do not establish that each primitive has this particular neural head. Likewise, Noul can ask a review question without implying a separate, documented “Act/Escalate” module inside Jev.

Calibrated decisions need an appropriate objective

TypeSafe calls its training method Reinforcement Learning for Calibrated Decisions, or RLCD. Its stated objective is to make the probabilities useful for decisions: predictions made with similar confidence should be correct at a corresponding frequency [1]. That is a measurable requirement. For binary outcomes YY, ideal calibration means:

E[Y∣q(X)=p]=p\mathbb{E}[Y\mid q(X)=p]=p

For example, among many comparable cases assigned probability 0.8, the event should occur about 80% of the time. A standard diagnostic for categorical predictions is the Brier score:

BS⁡=1N∑i=1N∑k=1K(pik−1[yi=k])2\operatorname{BS}=\frac{1}{N}\sum_{i=1}^{N}\sum_{k=1}^{K}\left(p_{ik}-\mathbf{1}[y_i=k]\right)^2

This penalizes probability assigned away from the observed outcome. It measures overall probabilistic prediction quality, including more than calibration alone; a reliability curve shows calibration more directly. This is an evaluation tool for the proposed system, not a disclosed formula for Jev’s proprietary training algorithm.

Type correctness and decision quality remain separate properties. A model can return a perfectly valid “continue” label and still make the wrong decision. Calibration therefore has to be checked on the actual process data, including new material lots and changed operating conditions.

The bridge to continual learning

In multi-head task-incremental learning, each task has its own classification head while the network shares a feature extractor. Labels from different tasks do not have to compete in a single softmax. A study by Lanzillotta, Meier, and Hofmann finds Neural Collapse structure within task heads, with variable alignment and scaling across tasks [4].

There is a useful analogy here. Continual learning separates heads by task; a manufacturing model can separate outputs by decision semantics. A quality class, a conductivity estimate, and a review probability carry different meanings and error costs. Each can have its own loss, validation set, and calibration procedure while drawing on common process features.

Shared features still create interference. Updating the encoder for one objective can move the representation used by every other head. That connects directly to my IGFA work: separate outputs help organize tasks, while retention under a changing representation remains a problem that the training procedure must address.

A process controller with several outputs

Battery manufacturing already lives in a world of multiple tasks and outputs: coating experiment design, calendering settings, impedance analysis, cycling prediction, and drift detection. The labels include pass/fail, risk levels, continuous properties, and decisions about whether to trust a prediction.

A shared encoder over process logs, measurements, experiment metadata, and surrogate-model outputs can support several readouts:

  • Choice: select a permitted next step, such as re-measuring a sample or requesting review.
  • Score: assess quality or risk against a defined rubric.
  • Regression: estimate porosity or conductivity with appropriate units and uncertainty.
  • Binary judgment: estimate whether a particular inconsistency or failure condition is present.

The resulting probabilities feed an explicit policy. With a loss L(a,y)L(a,y) for action aa when the true outcome is yy, a simple decision rule is:

a∗(x)=arg⁡min⁡a∈A∪{review}∑yp(y∣x)L(a,y)a^*(x)=\arg\min_{a\in\mathcal{A}\cup\{\mathrm{review}\}}\sum_y p(y\mid x)L(a,y)

This makes escalation depend on consequences. In an illustrative two-outcome case, suppose acting on a bad prediction costs 100 units and review costs 2. An 8% predicted failure probability gives an expected action loss of 8, so review wins; at 1%, the expected action loss is 1. The threshold comes from that loss model and assumes trustworthy probabilities, rather than from an arbitrary confidence percentage.

Historical data and simulations can train the heads, with later campaigns reserved for evaluation. The language model then handles experiment planning, documentation, and explanation around the decisions. This is the useful System One / System Two split: repeated judgments have a compact interface, and open-ended research retains a flexible reasoning tool.

Jeff makes the interface inspectable

The open-source Jeff 1 from Gestalt-Lab implements Jev-style Choice, Score, and Noul requests using a LoRA adapter on Qwen3-4B-Instruct-2507. It scores candidate labels from the language model’s token probabilities, using the first token when labels are distinguishable there and sequence scoring otherwise. This is a different implementation route to a typed interface [3].

The project publishes fact-checking evaluations and limitations, including confidently wrong judgments and weak performance when evidence is insufficient. Its results establish an inspectable starting point for that task; matching the API does not establish equivalent behavior on manufacturing data. The reusable part is the pattern: define the questions, supply representative labeled examples, fit or distill a model, and evaluate the probabilities and resulting decisions.

Jev and Jeff are reminders that every AI output does not need to be a sentence. Some outputs predict properties, some classify quality, some score risk, and some decide whether to ask for help. A common representation can support all of them. Give each decision a suitable output, then let the language model write the report.

Sources

  1. Almeida, D. Introducing System One Models & Jev. TypeSafe AI, 15 September 2026.
  2. TypeSafe AI. Official introduction and typed decision primitives.
  3. Gestalt-Lab. Jeff 1: implementation, evaluation, and training guide.
  4. Lanzillotta, G., Meier, D., and Hofmann, T. Heads Collapse, Features Stay: Why Replay Needs Big Buffers, 2026.
  5. Störk, J. The Geometry of Forgetting in Continual Learning and the calendering U-shape.

What should I call you?

Choose a display name for your comments. No email or account signup.

Your commenting identity

Use at least 12 characters. You’ll need this passphrase to restore the file.