Instrument theme
Jev comparison · Published evidence

Faster decisions.
Show the trade-off.

TypeSafe’s workflow evaluation puts Jev at the low-cost, low-latency end of the comparison. Quality still varies by task. Read the three measures together before choosing a model.

The Little Builder examining three different stacks of computing blocks with a magnifying lens
Evaluate your workload

Provider-reported data · Checked 18 September 2026

READ THIS FIRST

A workflow comparison, with a defined scope.

The charts reproduce TypeSafe’s rounded, four-workflow mean values for the workflow configurations. The tasks cover security incidents, agent trace observability, invoices and customer service. The score measures agreement with model-generated reference decisions, not independently labelled ground truth.

TypeSafe says the reference combines high-reasoning outputs from GPT-6 Astra and Claude Fable 5.1; compared models use their provider’s default reasoning setting. Display labels below are preserved exactly from the plot. Pinned model IDs and the aggregate sample count are not exposed in that overview, so this snapshot is not a fully reproducible independent benchmark. Source and methodology.

COST · TIME · AGREEMENT
Time per workflow

Seconds · lower is better · Linear scale from 0 to 90.0

Jev
0.4
terra
10.1
luna
12.9
haiku 4.5
12.5
sol
23.3
opus 5
37.8
DS v4 flash
51.9
sonnet 5
78.1
DS v4 pro
86.5

Equal-weight mean of four workflows. Includes the provider’s workflow harness; not token-generation speed.

Source: TypeSafe workflow evals. Checked 18 September 2026. Provider-reported, not independently measured by QQuantum.ai.

Cost per workflow

USD · lower is better · Linear scale from 0 to $0.1800

Jev
$0.0004
terra
$0.0304
luna
$0.0033
haiku 4.5
$0.0195
sol
$0.0836
opus 5
$0.1761
DS v4 flash
$0.0059
sonnet 5
$0.1174
DS v4 pro
$0.0413

Published rounded mean costs. The Jev bar is very small on this linear scale; use the printed value.

Source: TypeSafe workflow evals. Checked 18 September 2026. Provider-reported, not independently measured by QQuantum.ai.

Reference agreement

Percent · higher is better · Linear scale from 0 to 100.0

Jev
67.8
terra
67.9
luna
66.8
haiku 4.5
53.6
sol
74.1
opus 5
73.1
DS v4 flash
64.4
sonnet 5
67.8
DS v4 pro
65.5

Called accuracy by the provider. Agreement with reference-model decisions is not a universal quality score.

Source: TypeSafe workflow evals. Checked 18 September 2026. Provider-reported, not independently measured by QQuantum.ai.

WHAT FOLLOWS

A candidate for the decision layer.

The useful inference is architectural: test whether a specialised decision model can handle a narrow, frequent judgement while a larger model handles generation or harder reasoning. These results do not establish Jev as a better writer, coding agent or general problem solver.

Do not turn the rounded chart values into precise universal speedup claims. A short classification, a long document and a multi-stage workflow have different latency and cost profiles. Geographic distance and concurrency also change what users experience.

The next decision is whether Jev belongs in your model router, an agent checkpoint or a responsive application loop. Each needs its own evaluation.

MATCH THE INTERFACE TO THE JOB
Jev decision layer Generative model layer
Output needed A choice, score or probability within defined options An answer, explanation, draft or program
Tool-building role Select and assess actions that application code exposes Generate code or plan a broader sequence of work
Control Explicit options, thresholds and deterministic execution Validate generated output and constrain tool access
Evaluation question Did the decision produce the right downstream outcome? Did the generated artefact meet the task requirements?

Architecture guidance, not an additional benchmark. Both layers still require testing, permissions and failure handling.

YOUR EVALUATION

Compare complete systems on the same work.

Use representative inputs and a human-reviewed definition of success. Include difficult cases and ordinary traffic, not only examples that fit the model’s strengths. Keep a holdout set separate from the examples used to tune questions and thresholds.

Run the existing system and proposed integration with identical success criteria. Record accepted answers, escalation, latency percentiles and total cost, including checks and retries. Report results by workload category so a strong average cannot hide a weak critical path.

Document model identifiers, settings, location, concurrency, prompt versions, test date and sample sizes. Decide the quality floor before looking at the price. Our AI evaluation approach is designed around that decision.

QUESTIONS
QUESTIONS — 3

No. These charts reproduce TypeSafe’s published workflow comparison, checked on 18 September 2026. They are attributed provider results, not our own measurements.

Not automatically. The reference comes from other models and can be wrong or favour similar behaviour. Validate against the outcomes and reviewed labels that matter for your application.

We preserve the source plot’s display labels. The overview does not expose pinned model IDs, so we do not infer versions from those labels.

Keep reading

Your workload.
Your evidence.

We will define a fair comparison with your current system, then decide whether Jev earns a place in it.

CASE STUDIES

Shipped work.
Go and check it.

The work we can name, with the live site, our scope and the boundary made explicit. Select a project to see the evidence; each is a full case study, not a logo or a claim.

sonora.com
The Sonora homepage on desktop: a full-bleed dune landscape behind the headline “Transform Your Life with Sound”, with App Store and Google Play download buttons.
sonora.com — homepage, 1440×900 sonora.com →
Live Consumer wellness · Mobile + web

Sonora

Cognitive AI Ltd · 2026

A free sound-wellness app, described by its publisher as AI sound therapy that reads a short vocal sample at the start of a session and generates a soundscape for that moment. We designed and built the website and its backend, produced assets for the iOS and Android apps, and supported the application prototype.

Read the case study →

See every published project →

WHO WE HAVE BUILT FOR

Twenty-one years of applications, platforms and campaigns for names you know.

See all of our work →