A freight operations inbox receives a request for a quote, an update about a delayed shipment and a question about customs paperwork. Someone has to identify the work and send it to the right desk. A classifier can help with that first judgment. It still needs rules for incomplete requests, conflicting intents and the commitments the business allows it to make.
Laya, Jev and OpenAI Decisions address bounded judgments that software can consume. This guide explains the interface, compares the documented capabilities and follows a hypothetical freight request through review. We have not benchmarked these products against each other or deployed this example. The proposed workflow and sample outputs below are design illustrations.
What does an AI classifier actually return?
A classifier maps an input to categories you define. For an operations inbox, the input might be a message and relevant shipment context. The answer options might be quote request, shipment status, customs documentation and other. The application uses the selected category to assign work. Defining the options is part of the design: overlapping descriptions produce ambiguous tasks, while an incomplete list can force a poor choice.
Separate the questions when the answers serve different purposes. Request type determines the desk. Urgency determines its place in the queue. Whether an attachment contains the required fields determines whether an operator needs to ask for more information. Combining all three into one label makes the result harder to inspect or change.
Jev exposes Choice for selecting an option, Score for judging against an ordered rubric, and Noul for a yes/no probability. TypeSafe recommends narrow questions whose results are composed in code. Those are different output contracts, so choose the one that matches the application's next step. TypeSafe introduction
Laya, Jev and OpenAI Decisions: what we can compare
Jev's Choice interface accepts state and named questions with described options. It returns a selected option, a probability distribution and a confidence value. Its documentation supports an other option when the categories may not cover every input. A typed answer still requires application code to decide how it will be used. Jev Choice request and response
Laya publishes Apache 2.0 weights and a non-autoregressive encoder with a decision head. Its model card distinguishes English, multilingual and typed-decisions checkpoints. It also documents input and option-budget constraints, overconfidence, and substantial differences between base and fine-tuned results. Evaluate the exact checkpoint and runtime configuration you intend to operate; local hosting includes maintenance and compute costs. Laya model card and limitations
OpenAI's September 29 announcement describes Decisions API as Luna answering user-defined questions with finite predefined answers, using text or image context. It names classification, request routing and agent action selection as uses. The announcement says limited preview with broader release planned; it does not provide enough information to establish comparative accuracy, latency or cost. OpenAI Decisions announcement
| Product | Interface and operation | What to verify for your task |
|---|---|---|
| Jev | Hosted typed questions: Choice, Score and Noul. | Option descriptions, uncertainty handling, service terms and performance on your labelled cases. |
| Laya | Open weights; local inference; multiple checkpoints and typed answers. | Checkpoint, language, input length, option budget, tuning needs and hosting resources. |
| OpenAI Decisions | Announced finite answers from text or image context; limited preview described at launch. | Current access and API contract, supported outputs, pricing and measured task performance. |
For a first trial, choose the candidate you can access and operate under your data requirements. Local processing may matter when approved hosting boundaries prevent sending messages to an external service. A hosted API may reduce infrastructure work. Neither choice establishes better classification. A defensible comparison gives each candidate the same task definitions and authorized evidence, then inspects its errors.
Worked example: a freight request with two intents
Consider this invented message: 'Please quote collection in Miami tomorrow for delivery in New York. There are four pallets. Can your team also check the customs paperwork for our separate import shipment?' It contains a quote request and a second job. Its apparent urgency does not establish cargo readiness, equipment availability or a service commitment.
In a proposed design, code would attach a message identifier and only the context the classifier needs. It would keep account credentials and action tools out of the classification request. The following application-owned contract makes the labels explicit. It is illustrative JSON, not a universal API request for all three products.
{
"message_id": "example-001",
"request_type": {
"quote": "Ask for a price or available transport service",
"status": "Ask about an existing shipment",
"customs": "Ask for customs-document assistance",
"other": "None of these descriptions fits"
},
"additional_questions": [
"Does this message contain more than one request?",
"Is information needed before an operator can quote?"
]
}A single request-type question can miss the separate customs job even if it selects quote correctly. In this design, a second question flags multiple requests. The operator then splits the work, associates the customs question with the correct shipment, and asks for missing quote inputs such as pallet dimensions and weight. The review route exists because of the message's content, not just a low-confidence score.
{
"message_id": "example-001",
"request_type": "quote",
"multiple_requests": true,
"missing_information": true,
"next_step": "operator_review",
"permitted_external_action": "none"
}After review, the system could create separate internal queue items and preserve their common message reference. The quote workflow would use verified cargo details, current carrier data and the company's approval rules. Booking transport or issuing a price requires that later workflow's checks. The classifier's label authorizes neither.
Confidence, probability and accuracy are different
A returned probability describes the model's estimate for an answer. Confidence may be a separate statistic summarizing the distribution. Accuracy is the fraction of labelled cases the model answers correctly in an evaluation. None of those numbers substitutes for checking whether the requested action is allowed.
TypeSafe documents Choice confidence as a function of the top probability and number of options. With three options and a top probability of 0.6, its formula gives confidence 0.4. The two values describe the same prediction differently. Noul has no separate confidence field. Do not use a provider's confidence value as if it were the observed probability that your workflow will succeed. TypeSafe confidence definitions
To inspect calibration, group held-out predictions by their reported probability and compare each group with observed correctness. A group averaging 0.8 probability but getting only half its labels right would be overconfident on those cases. This is an invented illustration, not a finding about any named model. Provider-specific confidence statistics should remain identifiable when you normalize results; they are not interchangeable.
Choose review thresholds from the cost of an error and the evidence on your task. A queue assignment can usually be corrected. A shipment commitment may create costs that an internal assignment does not. High confidence can still accompany missing facts or the wrong label, so completeness checks and permission rules remain separate.
Evaluate the errors that change the operation
Build a labelled set from messages your team is authorized to use. Include routine cases, multiple requests, attachments, uncommon languages, missing facts and categories outside the proposed list. Ask operators to resolve disagreements about the labels before treating their decisions as ground truth. Keep evaluation examples separate from examples used to tune the model or question descriptions.
| Measure | How to inspect it | Operating consequence |
|---|---|---|
| Wrong routing | A confusion matrix shows the actual category against the predicted one. | Work reaches the wrong desk or requires reassignment. |
| Missed urgent requests | Count urgent cases the system failed to flag, alongside unnecessary alerts. | Deadline risk or an overloaded urgent queue. |
| Automatic coverage and errors | Measure the fraction assigned without review and the errors within that subset. | More automation is useful only if the released work remains acceptable. |
| Review workload | Count referred messages and measure the operator time to resolve them. | An accurate classifier can still move too much work into review. |
| Latency and cost | Measure the whole path, including preprocessing, API or hosting, retries and review. | Published inference speed alone does not describe the operating cost. |
Include a rules-based baseline. A structured request form or an existing shipment reference may already determine the desk without a model. Use the classifier for the cases where interpretation adds value. Compare changes on the same held-out set and examine individual misses, rather than picking a model from one aggregate accuracy number.
For this workflow, a release rule should combine evidence about routing quality with the team's available review capacity. A model that improves classification but overwhelms the review desk may not improve the operation. Record the acceptance criteria before testing so a promising demo does not quietly change the definition of success.
Where to start, and when to keep an operator
Start with one recurring, recoverable decision such as internal queue assignment. Run the proposal beside the current process without allowing it to make commitments. Keep the input reference, label definitions, model or checkpoint version, returned answer, review decision and actual downstream result. That record makes regressions and model changes inspectable.
Keep an operator involved when categories overlap, facts are missing, the input is outside the tested scope, or the action needs approval. If the task requires a long investigation, split out the short classification judgment and leave the investigation in a separate workflow. A bounded answer is most useful when the surrounding application makes clear what follows it and what happens when it is wrong.



