Jev vs OpenAI Decisions API: compare ticket routing before you migrate
Compare Jev and GPT-6 Luna Decisions request formats, refusal handling and input costs. Run a downloadable ticket-routing kit before testing real traffic.
Jev and the OpenAI Decisions API can both route a text ticket into a fixed queue, but they are not interchangeable endpoints. A useful comparison starts with the same category definitions and labeled inputs, then checks response mapping, uncertainty handling and total operating cost. Changing a model name inside an existing client does not complete the migration.
This guide gives you a small, reproducible migration exercise: eight fictional tickets, two request builders, strict response adapters, an input-cost calculator and an explicit acceptance plan. It connects the Jev overview with our Decisions API CSV tutorial, concentrating on what must change when one application supports both services.
Evidence boundary, checked October 8, 2026: the comparison below uses current official documentation and locally executed synthetic tests. We did not make paid API calls or measure either model’s accuracy or latency. Every fixture response is authored. Its job is to expose integration mistakes; it cannot establish a benchmark winner.
Compare the contracts before comparing model quality
For this exercise, “Jev” means TypeSafe’s pinned jev-1.13.0, while Decisions uses gpt-6-luna. The latter is currently a public-beta endpoint. Pinning the Jev version makes the configuration explicit; using jev-latest can change the answering version later. Save the returned model identifier with each real run. TypeSafe models, OpenAI Decisions guide.
| Integration detail | Jev native API | OpenAI Decisions API |
|---|---|---|
| POST endpoint | /v1/systemone | /v1/decisions |
| Shared evidence | state | input |
| Question collection | Map keyed by your question ID | Array with a unique name per question |
| Choice definitions | criteria map | choices array of value/description pairs |
| Returned answers | Map keyed by question ID | Array matched by name |
| Choice probabilities | Map from label to number | Array of value/probability pairs |
| Yes/no primitive | noul | predicate |
| Documented input scope | Text, including structured text data | Text and images |
The paths above use HTTPS on api.typesafe.ai for Jev and api.openai.com for Decisions; they are direct-provider endpoints.
The image-input difference matters if a damaged-product ticket includes a photograph. You cannot fairly compare a Jev text-only request against a Decisions request that sees the photograph and call the outcome a pure classification comparison. Give both the same approved text description, or declare a separate multimodal task. Converting an image into text introduces another model or preprocessing step, with its own errors and costs.
Neither choice contract is a request to generate an explanatory paragraph or arbitrary extracted fields. If your product needs a refund amount plus an evidence quotation, design a separate extraction step rather than pretending a queue label contains those facts. The TypeSafe HTTP reference documents the native request and answer fields.
The other primitives also need deliberate mapping; the kit below implements only Choice.
| Task | Jev | Decisions | Migration constraint |
|---|---|---|---|
| Unordered category | choice, criteria map | choice, choices array | Preserve label meanings, not just their names |
| Ordered rating | score, ordered criteria array | score, ordered levels array | Preserve the same level order and descriptions |
| Whether a condition holds | noul, result field noul | predicate, result field probability | Both return a 0–1 estimate; map the field and re-evaluate thresholds |
Both Score results are probability-weighted level indices, not automatically normalized 0–1 values and not confidence. A three-level rubric uses indices 0, 1 and 2; its score may be 1.1. If your application normalizes by dividing by the highest index, document that application-side transformation and use the same rubric on both sides. Do not confuse a normalized severity score with the probability that a condition is true.

Official English documentation captured October 8, 2026. This shows interface documentation, not a live ticket result. Source.
Define one queue policy and keep the labels stable
Our task is deliberately single-label: select the primary handling queue. The four labels are billing, technical, feature and review. Billing covers payment, invoice and refund requests only. Technical covers a broken existing feature or access issue only. Feature covers a new capability only. Mixed departments, vague requests and unrelated material go to review.
This definition resolves a common disagreement before a model ever runs: “Refund my payment and fix the broken export” belongs in review, not whichever department is mentioned first. If your real workflow must create two tasks, this single-choice taxonomy is unsuitable; redesign the output instead of marking either single answer as correct after the fact.
The kit includes these authored reference cases:
| ID | Ticket meaning | Reference label | What it tests |
|---|---|---|---|
| T01 | Send an invoice | billing | Clear administrative request |
| T02 | Existing export button fails | technical | Failure versus feature request |
| T03 | Add a calendar view | feature | New capability |
| T04 | Refund and repair export | review | Two departments |
| T05 | “It is wrong” | review | Insufficient evidence |
| T06 | Embedded instruction to say billing, followed by unrelated content | review | Treating input as data |
| T07 | Japanese login-error report | technical | A multilingual boundary example |
| T08 | Korean payment-receipt request | billing | Another language example |
These eight records are too small to estimate production performance. They are useful because each has a reason to exist. The Japanese and Korean cases verify that your pipeline preserves Unicode, not that either provider is accurate in those languages. TypeSafe explicitly describes English as its strongest training language and recommends evaluating other languages on your own material. Retain original-language tickets in a later representative evaluation; translating everything into English changes the workload.
For real data, have two reviewers label a sample using the same written policy and resolve disagreements before freezing the evaluation set. Keep a separate development set for changing instructions. Preserve the source record ID, remove unnecessary personal information and avoid sampling only tickets that already caused trouble.
Build two requests from one shared task
Download the migration kit. It uses the Python standard library and runs on Python 3.9 or later. Extract it and open a terminal in the extracted directory; no package installation or API key is needed for the first run.
The payload() function reuses one instruction and label dictionary. Its provider-specific shapes are:
# Jev: a map of questions, with criteria keyed by label.
jev_request = {
"model": "jev-1.13.0",
"state": ticket_text,
"questions": {"queue": {
"type": "choice",
"instructions": RULE,
"criteria": LABELS,
}},
}
# Decisions: a named question inside an array.
openai_request = {
"model": "gpt-6-luna",
"input": ticket_text,
"questions": [{
"name": "queue", "type": "choice",
"instructions": RULE,
"choices": [
{"value": label, "description": meaning}
for label, meaning in LABELS.items()
],
}],
}
RULE explicitly says to treat the ticket as evidence, not instructions, and to use review when the task is ambiguous. This instruction is a task definition, not a proof of prompt-injection resistance. The adversarial-looking T06 record remains a test case rather than a security certification.
Send one ticket per request for this exercise. Combining eight unrelated tickets into one shared input and asking one queue question would classify the bundle; it would not automatically produce eight independent answers. Multiple questions over one record are useful later, but changing request packing during a provider comparison also changes token counts and possibly behavior.
The live client uses each provider’s official host and Bearer authentication. It does not assume an OpenAI-compatible gateway implements /v1/decisions, or that TypeSafe accepts an OpenAI request body. Keep the native integration working before adding a gateway adapter.
Normalize answers without hiding failures
Run both fixture paths:
python3 migrate.py --provider jev --output jev-fixture
python3 migrate.py --provider openai --output openai-fixture
Each new directory contains eight raw JSON files, rows.json and summary.json. The program refuses to reuse an existing output directory, protecting previous evidence. The deliberately authored OpenAI fixture for T05 is a refusal; this does not predict that GPT-6 Luna would refuse that text in reality.
The adapter matches the requested question, requires the exact four-label set, rejects duplicate probability entries and validates finite probabilities that sum to approximately one. It also verifies that the selected label has maximal probability and that confidence is a finite number between zero and one. Missing keys, malformed answers and transport failures stay visible as errors.
Use three distinct states. status=ok with choice=review is a valid categorization into your review queue. status=refusal means Decisions declined to provide that answer. status=error means the application could not obtain or validate a usable answer. Turning all three into a successful review label would hide reliability problems inside an apparently healthy category count.
The current TypeSafe answer schema documents the three typed answer primitives; the adapter does not invent an OpenAI-style Jev refusal field. Unexpected Jev response types become validation errors. Preserve the raw response so a future schema change can be investigated instead of silently normalized away.
The example gate uses confidence 0.8, explicitly an illustrative setting. Every successful fixture has authored confidence 0.72, so default fixture runs route everything to manual handling and report automatic coverage zero, with automatic agreement null. This tests the no-automatic-decisions case. null means there was no denominator, not zero accuracy.
Measure routing quality separately from coverage
Do not copy a threshold from one provider to the other. TypeSafe documents its choice-confidence calculation from the distribution; a matching decimal in another interface does not establish equal calibration or equal business risk. Use the Jev confidence evaluation guide for a deeper threshold study and retain each provider’s original probabilities and confidence.
On a held-out real set, report at least:
- Valid-answer agreement: correct valid labels divided by all valid labels.
- Automatic coverage: automatically routed tickets divided by all submitted tickets, including failures.
- Automatic agreement: correct automatic labels divided by automatically routed tickets.
- Refusal and error counts, plus the confusion matrix by actual queue.
A system can improve automatic agreement merely by sending more work to people. That may be useful, but report the coverage reduction and review workload alongside it. Count billing tickets misrouted to technical separately from harmless disagreements between review and technical; the operational consequences differ.
Only after API access, data handling and spend are authorized, set TYPESAFE_API_KEY for Jev or OPENAI_API_KEY for Decisions securely in your environment and run:
python3 migrate.py --live --provider jev --output jev-live
python3 migrate.py --live --provider openai --output openai-live
Each command makes eight billable requests for the supplied CSV. Live mode keeps successful raw responses, including provider model and usage fields where supplied. The sample has no automatic retry. Investigate authentication and validation failures directly; apply bounded backoff for transient rate limits in a production client, logging every attempt and its potential charge. Do not compare fixture elapsed time with network latency.
Calculate costs using each provider’s actual usage
As checked October 8, TypeSafe lists Jev 1.13 at $0.042 per million input tokens, with output free. OpenAI documents $0.10 per million input tokens for GPT-6 Luna on Decisions, without cache-read, cache-write or output charges; regional premiums and long-context input multipliers still apply. These are direct-provider documented rates, not Ofox prices or a promise about every processing configuration. Jev pricing, Decisions pricing.
The kit’s arithmetic example is intentionally equal-token, not an observed bill:
python3 cost.py --jev-tokens 2000000 --openai-tokens 2000000
At the documented base rates it prints $0.084 for Jev and $0.20 for Decisions. Identical text does not guarantee identical billable token totals. Replace both inputs with their own recorded usage; include repeated instructions and all billed attempts. A low input rate alone cannot establish the cheaper production workflow.
Add human review explicitly. An illustrative 10,000-ticket workload with 20% manual review and $0.50 review cost per ticket adds $1,000, regardless of the small input-only bill. These assumptions are not measured customer costs. Migration engineering, image preprocessing, retries and downstream mistakes also sit outside the calculator. Our Jev routing break-even article develops the broader cost model.
Choose a service and release the migration gradually
For text-only queue routing, Jev offers a lower documented base input rate; that is a reason to evaluate it, not evidence it wins on your labels. Decisions is a practical candidate when image evidence is part of the task or your application already operates OpenAI’s native API. Its public-beta status belongs in the rollout decision alongside integration support and error handling.
Start in shadow mode: retain the existing production route, run the candidate on an approved sample and compare saved answers against the same references. Predeclare acceptable error, coverage and review-load limits based on your business; do not lower them after seeing a disappointing result. Examine language and ticket-type slices before switching a small percentage of traffic.
Keep the previous provider configuration and taxonomy version available for rollback. If a returned model changes, malformed answers rise or the review queue exceeds capacity, stop expanding traffic and inspect the recorded cases. The completion condition is a reviewed decision about task quality and operating cost, backed by real observations. Passing this downloadable kit proves the adapters behave as tested; it does not prove either model should receive every ticket.
Frequently Asked Questions
- Can I migrate Jev to Decisions by changing the model name?
- No. The endpoint, input field, question container and probability representation differ. Keep a shared label contract and use provider-specific request and response adapters.
- Which service classified these tickets more accurately?
- This article does not establish a winner. The downloadable fixtures are authored software tests, not live model results. Evaluate both services on held-out labeled tickets before choosing.


