A customer meets every condition for a refund and submits a request. A denial arrives immediately. The company's record reads as follows. This is a hypothetical example of an automated judgment shaping how a customer is treated.

Refund processing recordHypothetical example
Choice: REJECT
Choice probability: 0.87
Response: 200 OK

The denial has been sent, and the task is complete. The operations dashboard shows no sign of trouble. Yet the company's refund policy was not correctly applied to this customer.

TypeSafe's Jev points toward a future in which judgments like this are easier to embed throughout software. It takes unstructured information and returns a choice and probabilities that code can use directly.

If judgments can be invoked faster and at lower cost, there are more opportunities to handle customer requests automatically. There are also more places where companies need to check whether those judgments were appropriate.

1. Deploying AI judgment was expensive

Companies have long used machine learning in hiring, credit assessment, fraud detection and customer support. Its outputs affect who gets an interview, which transactions go through and whose requests are put on hold. Once a model's output is used in operations, the company has to manage the criteria behind it and the outcomes it produces.

One constraint on adoption has been the cost of building and running these systems. The cost here is not the compute used for a single prediction. It is the cost of building and maintaining a decision capability for a particular business task.

Developing a purpose-built model meant defining what counts as a correct answer, preparing data, training and validating the model, and connecting it to business systems. When the criteria or input conditions change in operation, the model has to be re-evaluated and revised. Adding a single judgment creates both a development project and ongoing operational work.

Building a decision capability: 01 Define correct answers, 02 Prepare data, 03 Train and validate, 04 Connect to business systems. When criteria or input conditions change: re-evaluate and revise.
▲ After launch, the work of adapting to changes continues.

Companies therefore start by automating the tasks that can justify this investment. High volumes or clear savings can justify investing in a decision, but building a separate model for every small check along the way is difficult. This cost never guaranteed fairness or safety. It did, however, act as an economic constraint that forced companies to choose how far to deploy machine judgment.

2. Jev could change the economics of AI judgment

Teams could already use classification results or structured outputs from LLMs to drive branches in code. What makes Jev interesting is its design emphasis on repeating these judgments quickly and cheaply.

Jev returns choices such as approval, rejection or further review, along with probabilities. A predefined choice can determine the next path the code takes, while the probabilities can help define the conditions for automatic handling and review. Judgment can then serve as a software primitive: a basic building block that can be added and combined at different stages of an application.

Instead of developing a purpose-built model, companies can try automating judgments by adding calls to Jev to an existing application. The starting point for considering a new use becomes a concrete question:

  1. Should we approve this refund?
  2. Does this transaction need further review?
  3. Should we interview this candidate?

If this approach proves economical, even small judgments that could not justify a separate development investment may become candidates for automation. As the cost and effort of adding a decision capability fall, the range of tasks companies can consider automating grows.

Adding a question does not finish the work of designing a workflow. Companies still have to decide what to ask, which choices to offer, how to combine multiple judgments and under what conditions to act on them. TypeSafe itself explains that real workflows require domain design: questions need to be broken down and combined consistently.

The more judgments companies automate, the more places they need to check whether those judgments can be trusted with real work.

Silver metal cans moving along a conveyor.

3. Failures can accumulate while operations look normal

As judgments become easier to add, AI may make more decisions within a single workflow. One refund request can involve classifying the request, interpreting the reason, checking whether enough information has been supplied, and applying payout conditions and exceptions. Each step uses different criteria and leads to different follow-up actions.

Repeating good judgment cheaply is real progress. But lower costs can also make it easier to repeat flawed judgments at scale. The following hypothetical hiring log records applicants' evaluation scores and outcomes.

Model choices and probabilities (hypothetical records): candidate #1042 REJECT at 0.87, #1827 REJECT at 0.91, #2184 PASS at 0.78, #3921 REJECT at 0.84
▲ Hypothetical hiring records. Choices and probabilities alone do not establish whether a decision was appropriate.

The model’s choices and their assigned probabilities are visible, but these records alone do not reveal the basis for those choices or how the probabilities were derived. They also do not show whether applicants’ qualifications were interpreted correctly or hiring criteria were applied properly.

Suppose that, after the system has processed 200,000 applications over three months, the company discovers that one group's rejection rate has been consistently higher. That difference alone does not establish an error or discrimination. It does, however, give the company a reason to investigate which conditions affected the outcomes. Even if gender or race was never used as a direct input, the company can examine whether information such as career gaps, schools or addresses served as proxies.

Problems do not always appear as differences between groups. Expressing the same information differently might increase false positives. Irrelevant context might flip decisions near a threshold. A business policy might change while judgments continue to follow the old criteria. Throughout all of this, system metrics can look normal:

System health (API errors 0, Schema failures 0, Uptime 100%, API latency normal) compared side by side with what still needs checking (were qualifications interpreted correctly, was screening consistent with business criteria, was the updated policy applied).
▲ A hypothetical situation from the article. System health and the appropriateness of judgments each need to be checked.

That is because the system successfully received the request and returned a response in the expected format. These metrics do not tell us whether the judgment correctly applied the business rules. If all that remains is a short decision value like REJECT, there are few clues for investigating what went wrong.

Every request may complete normally, yet the effects of those judgments accumulate inside the organization. Individual decisions need to be checked against business rules, while aggregate outcomes need to be examined for the conditions under which problems recur and how widely they occur. Even when a problem is found, the same error may recur if the decision criteria and the way decisions trigger actions remain unchanged. Preventing that requires controls that use what monitoring reveals to adjust the decision criteria and the scope of automatic action.

4. A model inventory may not be enough

One starting point for managing AI is an inventory of the models a company uses: which vendor and version, who approved the model and which benchmarks it passed.

But if systems like Jev make judgment easier to add throughout software, a model inventory may not be enough. The same model might be used to classify requests, check refund conditions and make a final rejection decision, yet each of those uses has different criteria and permitted actions. What appears as one model in the inventory becomes several places where the company needs to check a judgment. The vendor, version and approval status alone do not reveal where business rules may be applied incorrectly.

The same model (vendor, version, approval) branching into three decision points — request classification, refund eligibility checks, final rejection — each with its own criteria and permitted actions.
▲ Even with the same model, each decision point has its own criteria and permitted actions.

Knowing the ten models your company uses is different from knowing where those models make consequential decisions and which criteria they apply. That is why governance needs to extend from the model inventory to the judgments and actions themselves. For every decision, the company should be able to tell at least:

  • what input it saw and what judgment it made
  • which policy and threshold applied
  • whether that kind of decision was allowed to trigger an action automatically or required human review
  • what action actually followed
  • whether problems invisible in individual results are showing up in the aggregate

Whether a model is able to make a judgment and whether that judgment should be allowed to trigger a real action are two different questions.

If Jev is pointing in the right direction, the answer cannot be to make machine judgment expensive again. Good judgment getting cheaper and spreading faster is real technical progress. The trouble starts when the cost of making judgments falls while verifying and controlling them remains tied to people working by hand. Having humans re-read every decision an AI makes cannot keep up with that growth.

What is needed is a control layer that scales with judgment. It should check at execution time whether a judgment may turn into a real action, and let the company reconstruct later which evidence and policy allowed it. It should also show whether decisions that each look fine are piling up in the wrong direction overall. When a problem is confirmed, the company should be able to narrow the scope of automatic action based on that judgment, revise the criteria being applied or route decisions to human review, then evaluate whether the outcomes improve under the same conditions.

As the cost of AI judgment falls, controls and monitoring must scale with it.

We call the layer needed to solve this problem Decision Trust Infrastructure.