---
title: How to evaluate a decision model before you trust it
description: >-
  Decision models answer with a choice and a confidence. We tested Jev and open
  models on Airflow retries: how to compare them fairly before you trust one.
date: 2026-10-08T00:00:00.000Z
authors:
  - author: src/content/people/kaxil-naik.md
tags: []
categories:
  - ai
canonical_url: 'https://www.astronomer.io/blog/how-to-evaluate-a-decision-model/'
---
## Two and a half weeks after Jev

On 15 September [TypeSafe](https://typesafe.ai/) launched [Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev), a model that answers with a decision and a confidence score rather than text. By 1 October, several open models had appeared, including Cloudflare's [Clef](https://blog.cloudflare.com/clef-decision-models/) and AWS's [Strands Decider](https://github.com/strands-labs/strands-decider), and OpenAI had previewed a Decisions API (dates are in the table at the end).

[Apache Airflow](https://airflow.apache.org/) runs data pipelines as tasks, and a task that fails can be retried up to a number of times its author sets. A [retry policy](https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/tasks.html#retry-policies) lets Airflow decide failure by failure instead, and Airflow's Common AI provider package includes one that sends the exception to a decision model, then uses the category and confidence it returns to decide whether, and when, to retry. Jev was the first model we wired in. Exception text routinely carries hostnames, table and column names, and sometimes values from the rows being processed, which can be personal data. Many companies, especially in banking and healthcare, have policies that keep that kind of data away from models hosted outside their network. So in late September we tested whether a small open model that can run inside your own network could make the same call. We added [Kev](https://github.com/jaredpalmer/kev), Clef and Strands Decider in early October, after our first round of tests.

The number we cared about most was how often a model stops a task (fails it at once instead of retrying) when it would have passed on a retry. We count it two ways: as a share of the model's stops, which says how far you can trust a stop, and as a share of the recoverable errors the model was asked about, which says how often it blocks a recovery.

The short version: [decider-0.8b](https://github.com/Mapika/decider) came closer to Jev than we expected. At the same number of stops, the small open model's decisions were competitive with the hosted model's: the differences in wrong stops were small, and our sample didn't establish a reliable advantage for either. It ran on four CPU cores, though it answered far more slowly than Jev when five tasks failed at once. Neither is ready to stop tasks unwatched: when we let them stop a task only at a confidence score of 0.8 or higher, both stopped a fifth or more of the failures that, by our labels, a retry would have fixed. The rest of this post is how we measured that, and how to measure it on yours.

## Why a retry decision is hard

Every Airflow task that fails with retries left faces the same choice: try again, or stop. With `retries=5` and a 20-minute delay, a Snowflake login with the wrong password fails six times over more than an hour and a half while somebody waits for a table.

The two ways to be wrong cost different amounts. A wrong retry costs a delay and one more run, plus any side effect that isn't safe to repeat. A wrong stop fails a task that would have passed and gets somebody paged. So before looking at any model we decided that stopping should need more confidence than retrying, and that every answer the model isn't sure about should fall back to the task's normal retries.

The exception class alone often can't tell you whether to retry or stop. A JSON parsing error like this one looks like a reason to fix the input rather than retry:

```
json.decoder.JSONDecodeError: Expecting value: line 1 column 16385 (char 16384)
```

Jev and decider-0.8b, the open model we tested most, both said "bad data, stop", with confidence scores of 0.99 and 0.97. But char 16384 is a power-of-two boundary, which points at a truncated read. The task was reading its result from a helper container over a connection that sometimes cut off there, and the person who filed the issue wrote: "This occurs intermittently and works on next retry."

![The exception line with both models' verdicts (data, stop: Jev 0.99, decider-0.8b 0.97) above the issue comment that says it works on the next retry.](/images/posts/2026/how-to-evaluate-a-decision-model/img1_line_vs_thread.png)

## What we ask the model

A decision model takes three things: a description of the situation, a question, and a fixed list of answers, each with a short description. It returns one answer and a confidence. This is the request our evaluation sent Jev for the error above, which comes from our Airflow tuning set:

```json
POST /v1/systemone
{
  "model": "jev-latest",
  "state": "Classify this error from a data pipeline task (attempt 1 of 1):\n\njson.decoder.JSONDecodeError: Expecting value: line 1 column 16385 (char 16384)",
  "questions": {
    "response": {
      "type": "choice",
      "instructions": "You are an error classifier for a data pipeline system. Given an error message from a failed task, pick the single category that best describes it. Each category's description says what it covers.",
      "criteria": {
        "rate_limit": "API throttling or a quota exceeded.",
        "auth": "Credentials invalid, expired, or missing permissions.",
        "network": "Transient connectivity issue: connection reset, DNS, TLS handshake.",
        "data": "Schema validation, type mismatch, or bad input data.",
        "resource": "Resource not found or unavailable, such as a missing table or bucket.",
        "transient": "Temporary issue likely to resolve on its own.",
        "permanent": "Problem that will not resolve without a code or configuration change."
      }
    }
  }
}
```

Airflow's own policy sends the same kind of request, built from your categories, with the exception class and message and the task's real attempt number; our evaluation sent "attempt 1 of 1" for every error. The scored request also had no "unclear" option, which the appendix policy adds and we haven't scored.

These are the parts of Jev's reply our code read, with the values it saved:

```json
{
  "model": "jev-1.13.0",
  "answers": {
    "response": {
      "choice": "data",
      "confidence": 0.99,
      "probabilities": {"data": 0.99, "permanent": 0.01, "rate_limit": 0.0, "auth": 0.0,
                        "network": 0.0, "resource": 0.0, "transient": 0.0}
    }
  }
}
```

`jev-latest` resolved to jev-1.13.0. The policy compares `confidence` with the threshold; `probabilities` is how the model spread its answer across all seven options. The model only picks a category; whether that category means stop is the policy's choice. Our evaluation treats `auth`, `data`, `resource` and `permanent` as reasons to stop, and the other three as reasons to retry. `data` at 0.99 is a stop above the 0.8 threshold, so on its own this answer would fail the task on its first attempt, and the retry that would have worked never runs. That's why our evaluation ignores a model's answer on decode and parse errors and leaves them to the task's normal retries, and why the appendix policy retries them before asking a model.

## Building a test set from GitHub issues

We didn't find a public benchmark for "did retrying help?", so we built our own from GitHub issues: 400 errors from 35 projects data pipelines lean on, such as boto3, the Snowflake connector, psycopg and requests. Issue threads are useful because people say what happened next. We used a separate set of Airflow errors to choose the confidence threshold and a rule for decode errors, then evaluated those choices on the 400 errors described here. Neither set is public.

We used independent LLM-derived labels: raters separate from the models under test and blind to their answers read each thread and decided what a retry policy should have done (stop, retry, or can't tell). On the second set, three raters from two model families agreed on 176 of 200 errors. We took the majority on the rest and marked the errors they couldn't settle as disputed rather than forcing an answer. A stop counts as wrong when the labels say a retry could have fixed the error, as the thread describes it.

We built two sets of 200. The first leans toward bugs that fail every time. We relabelled it after first scoring it, because the original labels came from a single LLM assessment of each error, and our instructions told it to choose retry when unsure, which records the safe action rather than what the thread shows. Like the original raters, the new ones didn't see the models' answers. At the stop counts we compared, the relabel cut each model's wrong stops by as many as six, comparable to the differences between the models.

Both sets come from issues in the same 35 projects. The first gives a model few chances to make the mistake we care about most: stopping a task that a retry could have saved. So for the second set we deliberately included more issues describing intermittent failures, the kind a retry could fix, to test whether models can tell when retrying might help. We call it the intermittent-heavy set. Neither set necessarily reflects the mix of errors in your own pipelines.

Here are three rows from the intermittent-heavy set, with each model's answer and what the policy would do with it:

| Error (GitHub issue) | Label | Jev | decider-0.8b |
|---|---|---|---|
| `sqlite3.IntegrityError: datatype mismatch` ([SQLAlchemy](https://github.com/sqlalchemy/sqlalchemy/issues/10382)) | stop | data, 1.00: stops | data, 0.99: stops |
| `ResponseError: max number of clients reached` ([redis-py](https://github.com/redis/redis-py/issues/719)) | retry | rate_limit, 0.86: retries | rate_limit, 0.94: retries |
| `duckdb.TransactionException: TransactionContext Error: Failed to commit: Input is invalid/unsupported GZIP stream` ([duckdb](https://github.com/duckdb/duckdb/issues/8420)) | retry | data, 0.98: stops (wrong) | data, 0.91: stops (wrong) |

All three raters agreed on each label. The third is a wrong stop: its reporter titled the issue "Transient issue", but nothing in the error message says so, and our decode-error rule doesn't match it. We picked these three to show what a right stop, a right retry and a wrong stop look like, not as a typical sample.

## Handle familiar errors before asking a model

Before asking a model, we checked each error against rules we wrote, mostly from library documentation. They looked for familiar error codes and messages:

- Retry throttling and temporary server errors.
- Stop only errors that are nearly always permanent, such as missing imports, invalid SQL and rejected credentials.
- Leave decode errors to the task's normal retries, because they can result from interrupted reads.

Errors that matched none of these went to the model. Across both sets the rules made 25 stops, and none went against the labels, before or after the relabel. That left the models 41 errors labelled recoverable: 18 in the first set and 23 in the second.

Treat rules like these as your baseline, and judge a model by what it adds on top.

## 0.8 doesn't mean the same thing to every model

Every decision model reports a confidence, but each on its own scale. On the intermittent-heavy set, the median confidence was 0.94 for Jev, 0.84 for decider-0.8b and 0.45 for Fastino's [GLiNER2.5-Decide](https://huggingface.co/fastino/GLiNER2.5-Decide). So the same 0.8 threshold means something different for each model.

It can also make a model look safer than it is. At 0.8, GLiNER2.5-Decide made only 10 stops and got just 1 wrong, the fewest wrong stops of the three. But six of those ten stops came from the rules, so it was mostly just deciding less.

To compare the models fairly, we let each make the same number of stops, the rules' stops plus that model's most confident ones, and counted how many were wrong. At 40 stops on the intermittent-heavy set, GLiNER2.5-Decide made 5 wrong stops, decider-0.8b 1, and Jev between 1 and 3, depending on how ties among its equal confidences are broken. Across all the stop counts we could compare on both sets, decider-0.8b and Jev stayed close: at most stop counts Jev was level or behind, and at some it was ahead by up to two. With only 18 and 23 recoverable errors per set, we couldn't establish a reliable difference.

![Wrong stops against the number of tasks stopped on the intermittent-heavy set, with rules first and then each model's most confident stops. The lines are never more than three errors apart, and they touch and cross.](/images/posts/2026/how-to-evaluate-a-decision-model/img7_v6.png)

We ran five newer models through the same evaluation with the same inputs, including category descriptions written before either set existed: Kev 0.8B and 4B, Cloudflare's Clef-flash (9B) and Clef (27B), and Strands Decider 2B. All five are Apache-2.0. At equal numbers of stops, every 20 from 40 to 100 on both sets (the tables are in the appendix), we found no clear evidence that another model performed better than decider-0.8b. One comparison favoured decider-0.8b over GLiNER2.5-Decide, but an isolated finding among 55 comparisons needs further confirmation.

0.8 is the threshold we actually ran, so this table shows what it caught and what it cost for each model. Both sets together, counting only the model's own stops (the rules add 25, none wrong):

| Model | Median confidence | Stops | Wrong stops (of 41 recoverable) | Right stops (of 222 that should stop) |
|---|---|---|---|---|
| decider-0.8b | 0.84 | 135 | 9 | 122 |
| Jev 1.13.0 | 0.94 | 147 | 15 | 128 |
| GLiNER2.5-Decide | 0.45 | 6 | 1 | 5 |
| Kev-0.8B | 0.38 | 27 | 0 | 27 |
| Kev-4B | 0.43 | 13 | 0 | 13 |
| Clef-flash | 0.83 | 133 | 11 | 118 |
| Clef | 0.77 | 112 | 5 | 105 |
| Strands Decider 2B | 0.55 | 63 | 2 | 61 |

Read the columns together: Kev and Strands make few wrong stops at 0.8 because fewer of their answers clear it, so they also stop far fewer of the errors that should stop. For decider-0.8b, 9 of its 135 stops were wrong (Wilson 95% interval 4% to 12%), and it stopped 9 of the 41 recoverable errors it was asked about (12% to 37%); counting the rules' decisions too, that's 9 of 160 stops and 9 of 55 recoverable errors. Of those 41 errors, Jev alone stopped 9 that decider-0.8b didn't, decider-0.8b alone stopped 3, and both stopped 6, not enough to call either one safer. Median confidence is over the intermittent-heavy set's 166 errors a running task would raise, and stops that are neither wrong nor right went to errors the raters couldn't settle.

The lesson: don't carry one threshold across models. Choose a threshold for each model on calibration data, and check it on errors you haven't used.

## Wording and what the message leaves out

The descriptions you give the categories are part of the setting. We reworded ours twice, trying to keep the meaning. That changed 5 to 14% of each model's own stop/no-stop decisions at 0.8, over all 200 errors in each set, for both Jev and decider-0.8b, and both rewordings cut decider-0.8b's stops by about a fifth. Even with the wording fixed, Jev's confidence moved by up to 0.14 between repeated calls, enough to flip 3 or 4 of 200 decisions near the threshold. Treat the descriptions and the threshold as one setting, and re-run your evaluation whenever you change either.

The message also leaves things out. Besides the DuckDB error above, a 404 the maintainers called a temporary server problem and an S3 Select read cut short under concurrent load were intermittent by their threads, but neither message says so, and both models stopped all three. Two of each model's wrong stops at 0.8, that 404 among them, came from the `resource` category, which our evaluation treated as a stop even though its description includes "unavailable". Tracebacks might help: an early test on our tuning set pointed that way, on too few recoverable errors to tell. A traceback also sends file paths and source lines to the model. An `include_traceback` option has [merged](https://github.com/apache/airflow/pull/74308) into Airflow's main branch.

## Running the policy in Airflow

There are three ways the model can fail to give a usable answer, and each should leave the task where it would have been without the policy:

- It doesn't know. The model can answer "unclear", and the policy then retries with the task's own delay.
- It isn't confident enough. Below the threshold, the task keeps its configured retry behaviour.
- It's too slow. Past the policy's timeout, 30 seconds by default, the task falls back to the same retry settings. In an earlier run of our demo against a cold server, three of five calls hit it, so set yours from the burst latency you measure.

The demo Dag (Airflow's name for a pipeline) used the policy in the appendix, a decode-error rule and then the model, with six constructed failures: two credential failures (a wrong password and a revoked token) that fail every time, and four that recover on retry. Neither model stopped any of the four that recover. Jev stopped both credential failures on the first attempt, and the whole run took 33 seconds. decider-0.8b also stopped the revoked token straight away, but it was less sure about the wrong password (0.67 to 0.70, under 0.8), so that task used all three attempts, five minutes apart, and the run took ten and a half minutes, most of it retry delay. A credential rule like the one in "Handle familiar errors before asking a model" would have stopped it on the first attempt. Every decision lands in the task log:

![Task log for snowflake_login on Jev: the Snowflake traceback, then "Retry policy decision action=fail reason=ClassifierRetryPolicy: category=auth confidence=0.99 threshold=0.80".](/images/posts/2026/how-to-evaluate-a-decision-model/shot_log_jev_snowflake.png)

If you try this on your own Dags, start in retry-only mode: set `retry=True` on every category. The model then only picks the delay, and the log still records the category it chose, so you can see which failures it would have stopped and what happened to them. A task that recovers shows that stopping it would have been wrong. The reverse is weaker: a task whose retries all failed shows only that they ran out within its allowed retries and the delays the model chose, not that a longer wait wouldn't have worked. Before you turn stopping on, decide the highest wrong-stop rate you'd accept and how many of your own failures you want to see first.

## Latency when many tasks fail at once

The policy runs on the worker when a task fails, so a slow answer keeps that worker slot busy until the model replies or the timeout hits. A call that times out falls back to the task's normal retries, and the model's choice of delay is lost.

We ran decider-0.8b's own server, which speaks Jev's API, on [Modal](https://modal.com) in a container with 4 vCPUs and 8 GB, and the Airflow workers called it over HTTPS. The demo rows come from the Airflow run described above, where five of its six failures reach the model.

| What we measured | Time |
|---|---|
| Jev, one decision in the demo | 0.26 to 0.38 s |
| decider-0.8b on 4 vCPUs, one request inside the container | about 1.6 s |
| decider-0.8b on 4 vCPUs, one call from Airflow over HTTPS | 2.5 to 4.5 s |
| decider-0.8b on 4 vCPUs, 5 tasks failing at once in the demo | 4 calls took 13 to 16 s, the fifth 2.5 to 5 s (two runs) |

A lone call looked manageable. Five tasks failing together pushed four of decider-0.8b's calls to about 13 to 16 seconds, and an outage can fail hundreds of tasks at once. In a separate burst of six requests sent inside the container, Strands Decider 2B answered all six in under 4 s each, while decider-0.8b's server took 9 s for five of the six; that's one burst each, and we don't know whether the difference comes from the server or the model. Hosting a model also means operating that service, patching and scaling included.

## What we'd recommend

Start in retry-only mode, whichever model you pick: at 0.8, both Jev and decider-0.8b stopped a fifth or more of the recoverable failures they saw. A benchmark like ours can shortlist models; your own logged decisions are what tell you whether to turn stopping on.

If your errors can't leave your network, decider-0.8b is the model we tested most, and none of the others was measurably better; size its server for your worst burst and keep it warm. Strands Decider answered our one burst faster, but its confidences run low, so it needs its own threshold, and we'd first test whether that advantage comes from the server or the model. If your errors can leave your network, Jev answered in 0.26 to 0.38 s in our demo, and there's no server for you to run.

Whichever you pick, handle errors with known retry behaviour using rules before asking the model (in Airflow, a small custom retry policy; see the appendix). To compare models on your own failures, log the exceptions and what happened to each task, then replay them through each candidate offline and compare at the same number of stops. If you build your own set, spend human checking on the few dozen errors labelled recoverable: they decide every comparison (41 in ours).

The errors we tested on come from bug reports rather than production, and most got stop labels. We don't know whether that mix resembles your failures. We didn't test whether you need a decision model at all rather than a general-purpose LLM behind Airflow's `LLMRetryPolicy`.

For Astronomer customers, the model gateway we mentioned in [our last post](https://www.astronomer.io/blog/what-jev-will-do-to-data-engineering/) will offer Jev under a zero data retention agreement. Deployments that can't send exception text outside their network would still need a local option.

Next we'd like to put OpenAI's Decisions API through the same evaluation once we have access, and repeat the comparison with tracebacks once that option is released. The appendix has an example Airflow configuration, and the [Common AI retry policy docs](https://airflow.apache.org/docs/apache-airflow-providers-common-ai/stable/retry_policies.html) cover every option. If you run it on your own failures, we'd like to see the errors it gets wrong.

## Release dates

| Date | What appeared |
|---|---|
| 18 Sep | [Laya](https://huggingface.co/convaiinnovations/laya), a 421M open decision model |
| 19 Sep | [decider](https://github.com/Mapika/decider), open models that answer on Jev's API |
| 20 Sep | Jared Palmer's [Kev](https://github.com/jaredpalmer/kev) |
| 22 Sep | vLLM merges a [Jev-style mode](https://github.com/vllm-project/vllm/pull/57250) for DiffusionGemma |
| 23 Sep | Fastino's [GLiNER2.5-Decide](https://huggingface.co/fastino/GLiNER2.5-Decide) |
| 28 Sep | [Ollama 0.35](https://github.com/ollama/ollama/releases/tag/v0.35.0) serves decision models over Jev's API |
| 29 Sep | OpenAI previews a [Decisions API](https://openai.com/index/devday-2026-recap/) at DevDay |
| 30 Sep to 1 Oct | Cloudflare's [Clef](https://blog.cloudflare.com/clef-decision-models/) and AWS's [Strands Decider](https://strandsagents.com/blog/introducing-strands-decider/) |

## Appendix: wrong stops at equal numbers of stops

These are total stops, rules included: the rules supply six on the intermittent-heavy set and 19 on the first. We picked this grid after seeing the results, so read it as a sensitivity check, not a planned test; below 40, about half or more of the first set's stops are the rules'. Where a model reports equal confidences at the cut, the range covers every order of the ties. Strands Decider at 100 stops on the intermittent-heavy set has no bootstrap interval, because in about a fifth of resamples too few of its answers picked a stop category to reach 100 stops. At the 0.8 threshold, an exact McNemar test on the 41 recoverable errors gives p = 0.15 for the difference between Jev and decider-0.8b.

| Intermittent-heavy set: wrong stops at | 40 | 60 | 80 | 100 |
|---|---|---|---|---|
| decider-0.8b | 1 | 2 | 6 | 10 |
| Jev 1.13.0 | 1 to 3 | 5 | 7 | 10 |
| GLiNER2.5-Decide | 5 | 5 | 6 | 10 |
| Kev-0.8B | 1 | 5 | 9 | 14 |
| Kev-4B | 1 | 3 | 7 | 9 |
| Clef-flash | 1 | 4 | 8 | 8 |
| Clef | 1 | 3 | 5 | 9 |
| Strands Decider 2B | 2 | 5 | 7 | 12 |

| First set: wrong stops at | 40 | 60 | 80 | 100 |
|---|---|---|---|---|
| decider-0.8b | 1 | 1 | 3 | 7 |
| Jev 1.13.0 | 0 to 3 | 5 | 6 | 8 |
| GLiNER2.5-Decide | 0 | 1 | 2 | 5 |
| Kev-0.8B | 0 | 0 | 3 | 6 |
| Kev-4B | 1 | 2 | 5 | 6 |
| Clef-flash | 0 | 4 | 6 | 6 |
| Clef | 2 | 2 | 3 | 4 |
| Strands Decider 2B | 0 | 1 | 3 | 5 |

In an earlier round we also tried GLiClass and Laya, which weren't useful for this policy as we ran them, and decider-35b and DiffusionGemma-26B, which showed no overall advantage worth their larger serving footprint.

## Appendix: Airflow configuration

This is not the policy we scored above. It has one rule, which matches decode errors by exception type and retries them, and then the model. It leaves out the message-based rules from "Handle familiar errors before asking a model": Airflow's `RetryRule` matches exception types, not message text, so those rules need a custom retry policy. Its category descriptions and thresholds also differ from the ones we scored, so re-run the evaluation before you rely on it.

Swapping models is one connection:

| | Jev | Self-hosted decider-0.8b |
|---|---|---|
| Connection Id | `jev_default` | `decider_local` |
| Connection Type | Pydantic AI | Pydantic AI |
| Host | empty (TypeSafe's endpoint) | `https://<your-decider-server>` |
| Password | your TypeSafe API key | the key your server or proxy checks |
| Extra | `{"model": "typesafe:jev-1.13.0"}` | `{"model": "system-one:decider-0.8b"}` |

The policy takes `llm_conn_id="jev_default"` or `llm_conn_id="decider_local"` and nothing else in the Dag changes. The `system-one:` prefix reaches any decision model served over the System One API (the API Jev introduced), which decider, Kev and Strands Decider all speak, and needs pydantic-ai-slim 2.53 or later. Some of these servers reject a question without instructions or an option without a description; the retry policy always sends both. We ran the demo before `system-one:` existed, through pydantic-ai's TypeSafe client pointed at our server, and have since checked the `system-one:` connection against decider-0.8b.

The policy needs Common AI provider 0.10.0. `ChainRetryPolicy` is in Airflow 3.4, and Common Compat 1.20.0 provides it on 3.3; we ran the demo on Airflow's development branch.


```python
import json
from datetime import timedelta

from airflow.providers.common.ai.policies.retry import (
    ClassifierRetryPolicy,
    ErrorCategory,
)
from airflow.providers.common.compat.sdk import ChainRetryPolicy
from airflow.sdk import ExceptionRetryPolicy, RetryAction, RetryRule

CATEGORIES = {
    "rate_limit": ErrorCategory(
        "The service throttled the request or a quota ran out: "
        "HTTP 429, ThrottlingException, rate exceeded.",
        delay=timedelta(seconds=30),
    ),
    "network": ErrorCategory(
        "The connection failed: refused, reset or timed out, DNS or "
        "TLS errors, lost connection to a server.",
        delay=timedelta(seconds=10),
    ),
    "transient": ErrorCategory(
        "A temporary server-side condition that clears by itself: 5xx, "
        "deadlock, lock or statement timeout, warehouse busy or resuming.",
        delay=timedelta(seconds=20),
    ),
    "auth": ErrorCategory(
        "Credentials were rejected or a permission is missing: wrong "
        "password, revoked or expired grant, AccessDenied, 401, 403. "
        "Not a dropped connection.",
        retry=False,
    ),
    "data": ErrorCategory(
        "The task's own input is invalid and will be invalid on every "
        "run: schema mismatch, constraint violation, a value that cannot "
        "be parsed. Not a decode error on a network response or a "
        "stream, which is often a truncated read.",
        retry=False,
        min_confidence=0.9,
    ),
    "permanent": ErrorCategory(
        "A code or configuration bug that fails the same way on every "
        "attempt: import error, wrong argument, SQL syntax error, "
        "missing setting.",
        retry=False,
        min_confidence=0.9,
    ),
    "unclear": ErrorCategory(
        "The error does not say whether another attempt would succeed: "
        "a generic wrapper, an error raised while handling another "
        "error, or something that could be intermittent."
    ),
}

retry_policy = ChainRetryPolicy(
    [
        # Treat every decode error as possibly a truncated read and
        # retry it. This matches the exception type only; it can't tell
        # a cut-off stream from bad data.
        ExceptionRetryPolicy(
            rules=[
                RetryRule(
                    exception=[json.JSONDecodeError, UnicodeDecodeError],
                    action=RetryAction.RETRY,
                    retry_delay=timedelta(seconds=10),
                    reason="decode error, possibly a truncated read",
                )
            ]
        ),
        ClassifierRetryPolicy(
            llm_conn_id="jev_default",
            categories=CATEGORIES,
            min_confidence=0.8,
            timeout=30,  # the default; set it from your measured burst latency
        ),
    ]
)
```
