Logo

Decision models are here. Here's what happened when you put Jev inside a Dataiku agent.

September 30, 2026/Christiaan Burrett & Youssef Jouini

TypeSafe's Jev is one of a new class of decision models, alongside OpenAI's newly announced Decisions API. It returns a different kind of answer: a typed judgment with probabilities. Ask a yes/no question, choose between categories, or score a request against ordered criteria, and the result is ready for the next step in an application. For teams automating the small decisions inside business processes, that's an appealing proposition, especially at a few cents per thousand requests in our test.

But an enterprise workflow needs more than a promising model. It needs a way to connect that model to business rules, compare its behavior with alternatives, and decide where people should stay involved.

We brought Jev into Dataiku to do exactly that: build a reusable agent block, put it alongside LLM agents, and measure quality, response time, and cost on the same 200 supplier requests.

A new model, a native building block

Dataiku's Structured Visual Agents assemble a workflow from blocks. A plugin can extend that palette with a new capability, so we packaged Jev as a single Jev Question Set block.

Question Set block in a Structured Visual Agent

The Question Set block in a Structured Visual Agent: each judgment is a yes/no, choice, or score question written in plain language

A builder defines the judgments they need, chooses a type for each—yes/no, choice, or score—and writes the criteria in plain language. The block asks all the questions in a single call to Jev via Dataiku's LLM Mesh and stores the answers in the agent's state, ready for routing or other actions. 

The integration seamlessly brings Jev onto the visual agent canvas. Its judgments become part of a workflow that a team can inspect, test, and iterate on.

One business task, three agents

Our test case is supplier onboarding. Each request needs one of three outcomes: the standard onboarding queue, human review for a policy exception or sensitive access, or a request for missing details.

Each agent makes three judgments: is an exception being requested, what kind of request is this, and is there enough information to proceed? A native Routing block then applies the business policy. One agent uses Jev; the other two use GPT-5.6 Luna and GPT-5.6 Sol through Dataiku's LLM Mesh with a structured output schema.

The agent in the question set block

The agent: one Question Set block makes the three judgments, and a native Routing block applies the right policy

All three address the same task and routing intent. Their judgment interfaces differ: Jev supplies probabilities and scores, while the LLMs follow schema-based prompts. This comparison measures the three configured agents as a team would use them in a workflow.

The flow in the question set block

The Flow: 200 labeled requests run through all three agents

What putting them to the test revealed

We ran the three agents on 200 supplier requests, each labeled in advance with the right outcome. Many are deliberately tricky: the deciding detail is buried in ordinary procurement talk. Example: "Please raise a PO for Hartwell Logistics to move the archive boxes to the new storage unit on the 14th. Standard terms, quote attached. My director is on leave until the end of the month, so I'll approve it as both first and second approver to keep things moving." Everything looks routine until the last sentence. Jev missed it.

Jev 1.13

GPT-5.6 Luna

GPT-5.6 Sol

Routed as expected

157 / 200

169 / 200

174 / 200

Routed as expected in %

(79%)

(85%)

(87%)

Mean agent response time

0.3 s

3.0  s

2.2  s

Response time p95

0.4 s

4.7  s

4.0  s

Cost per 1000 requests

$0.03

$0.25

$4.35

Jev is about 10× faster and far cheaper. It answers in about 0.3 seconds, and 95% of its answers arrive in under half a second; the LLM agents take 2 to 3 seconds. Jev costs a ninth of Luna and under 1% of Sol.

The LLMs are more accurate. The hard requests were mostly hard for everyone: when Sol got one wrong, Jev usually did too. The gap comes from the requests only Jev missed, above all those that needed a person: Jev let about a quarter of them through without review (26%), against 6% for the LLMs. Jev did win on some vague requests, though. On 5 requests too thin to act on, such as "Can we set up Stellan Partners for some consulting? Thanks.", it asked for details while Sol sent them straight to the queue.

Jev's probabilities give the policy an extra dial. Most of the reviews Jev missed were close calls: it saw a possible exception but scored it below the 0.70 threshold. Lowering the threshold to 0.30 would have caught 10 more of them, at the cost of 4 unnecessary reviews. A team could go further and send only the borderline judgments to an LLM or a person, keeping Jev's speed and cost for the clear-cut cases. We haven't tested that hybrid yet.

Three lessons for adopting decision models

Own your model choice. Decision models arrived with a different interface than chat models; thanks to the flexibility of the LLM Mesh, builders can configure and run them alongside LLMs on the same task. With Dataiku, teams bring a new model onto the platform as soon as it launches, then decide per use case which one best fits their trade-off among quality, speed, cost, and where their data goes (Jev runs as a hosted API, while open-source alternatives are emerging).

Keep the business rule separate from the model. In each agent, the model only answers three questions; a native Routing block holds the policy. Swapping the model or moving a threshold never touches the rule, so choosing a cheaper model or a stricter policy is a configuration change that can be tested before it ships. Structured Visual Agents make that separation the default.

Evaluate against your own ground truth, in the Flow. The evaluation is an ordinary Dataiku Flow: a labeled dataset of requests, Prompt recipes that run each agent, Agent Evaluation that offers the comparison natively, and a scenario that reruns everything. Business and technical teams have one place to decide which approach fits and to reassess that decision when a model or requirement changes.

Decision models bring a fast, inexpensive typed-judgment primitive to the AI toolkit, and new ones keep arriving. Dataiku provides the place to turn them into business workflows: integrate them, compare them, compose them with explicit rules, and measure the decisions they produce against the data that matters.

Try it yourself. The plugin is open source on GitHub. Drop the Jev Question Set block into your own Structured Visual Agent.

Ready for AI success?