[ wearables · health ]
Health in real time.
Fast and affordable enough to decide on every sensor reading, not once a day.
- Context:
- reading out of range
- Decision:
- alert someone? yes
- Confidence:
- 0.93
- Action:
- notify a family member or professional
Our first decision model. You send a situation and the questions you want answered, with the possible options. Iryx picks the answer to each question and tells you how sure it is. It doesn't write text: it decides.
Measured on Iryx's own answer key on 2026-09-23, on short texts in Portuguese. Next to nine other systems: see where it does well and where it falls behind.
Three steps. You ask, Iryx decides and tells you how sure it is, and you choose when the system can act on its own.
[ 01 · you send ]
The context is the situation: a message, a record, a JSON. With it go the questions, each with its possible options. There are three kinds: pick one option, yes or no, a level on a scale.
I bought sneakers from your store last week and they came in the wrong size. [...] Can I exchange them by Friday?> questions
[ 02 · iryx decides ]
For each question, Iryx returns the chosen option, the confidence (from 0 to 1) and the spectrum: the chance of every option, not only the chosen one.
[ 03 · you set the threshold ]
You choose the minimum confidence for the system to act on its own. Anything below goes to a person. The band (high, medium, low) is the cut we recommend.
Example: the sneaker exchange from Try it now, a real test run on 2026-09-24 with a made-up message in Portuguese. The three kinds of question and the values are the ones from that test. The 0.80 threshold is only an example: you set it.
Glossary
A context goes in, decisions come out. Each one with the value, the confidence and the full spectrum. The band tells you when to check.
Pick a message:
POST /v1/decide< decisions
[ benchmarks · own answer key ]
Nearly the same accuracy as the largest models, in a fraction of the time per case. On math, it gets every one right. Ten systems, the same questions, our own answer key.
On the full main test, 82.2%, against 84.9% to 88.4% for the general-purpose models. Per case, ~90 ms; the general-purpose models, 802 to 4,747 ms, measured in batches: not the time of one request. The open Nimble 9B, 568 to 1,790 ms on a small home graphics card. The same cases for all, in the same order in every chart.
Accuracy on our own answer key. * Time includes the agent or command line that called the model, in batches: not the time of one request.
How to read it
Each square is one complete case, with ~3 decisions, at each system’s median time. It illustrates time per case; it isn’t a throughput measure. Iryx: 200 requests, 2026-09-23. GPT-5.5, GPT-5.6 Luna and Nimble 9B: 2026-09-30. The others: 300 cases, 2026-09-21.
Iryx, Jev 1.13 and Nimble 9B
Per request, through the API, with a new connection per request. Jev goes over the internet, with TLS; Iryx doesn’t: the gap between them isn’t only the model. On a kept-open connection, Iryx takes ~74 ms. Nimble 9B runs locally, through Ollama, on a small home graphics card: it doesn’t represent a suitable card.
* In batches
The six general-purpose models were called in batches, by an agent or from the command line. The time is the batch time divided by its cases, start-up included: it isn’t the time of one request. In a batch, the cases come out together, at the end.
| System | Math 33 | Main | Blind 1 100 | Blind 2 100 | |
|---|---|---|---|---|---|
| Text only 192 | Full 225 | ||||
| Iryx 1.0 | 100% | 79.2% | 82.2% | 86% | 73% |
| Jev 1.13 | 57.6% | 84.9% | 80.9% | 87% | 92% |
| Nimble 9B | 42.4% | 76.6% | 71.6% | 86% | 88% |
| Nimble 9B q4 | 42.4% | 72.9% | 68.4% | 86% | 87% |
| Claude Opus 5 | 93.9% | 87.5% | 88.4% | 96% | 97% |
| Claude Fable 5.1 | 87.9% | 88.5% | 88.4% | 96% | 93% |
| GPT-6 Astra | 90.9% | 86.5% | 87.1% | 95% | 96% |
| GPT-5.6 Sol | 87.9% | 87.5% | 87.6% | 96% | 95% |
| GPT-5.6 Luna | 87.9% | 84.4% | 84.9% | 86% | 91% |
| GPT-5.5 | 84.8% | 87.0% | 86.7% | 92% | 93% |
Accuracy on our own answer key; the number under each test is the question count. The full main test adds the 33 math questions to the 192 text ones. Answers from 2026-09-21 (math, main and blind 1) and 2026-09-22 (blind 2), measured on 2026-09-22 and 2026-09-23; for GPT-5.5, GPT-5.6 Luna and the two Nimble 9B, and for GPT-5.6 Sol on blind test 2, answered and measured on 2026-09-30.
| System | Time per case median | Cases | How it was called |
|---|---|---|---|
| Iryx 1.0 | 90 ms | 200 | API, not over the internet, one request at a time, new connection per request (on a kept-open connection: 74 ms) |
| Jev 1.13 | 412 ms | 300 | Vendor API, over the internet, new connection per request |
| Nimble 9B | 1,790 ms | — | local Ollama API (/v1/systemone), one request at a time, on a small home graphics card the model doesn’t fit on whole (part in CPU memory) |
| Nimble 9B q4 | 568 ms | — | same, in the compressed version (q4), which fits whole on the same card |
| Claude Opus 5 | 802 ms* | 300 in 6 batches | agent, batches of 50; batch time ÷ 50 |
| Claude Fable 5.1 | 1,515 ms* | 300 in 6 batches | agent, batches of 50; batch time ÷ 50 |
| GPT-6 Astra | 4,069 ms* | 300 in 30 batches | OpenAI command line, batches of 10, low reasoning; process time ÷ 10 |
| GPT-5.6 Sol | 4,747 ms* | 300 in 30 batches | OpenAI command line, batches of 10, low reasoning; process time ÷ 10 |
| GPT-5.6 Luna | 2,714 ms* | — | command line (Codex), batches of 10; batch time ÷ 10 |
| GPT-5.5 | 2,491 ms* | — | command line (Codex), batches of 10; batch time ÷ 10 |
A case is one request with ~3 decisions. Iryx measured on 2026-09-23; GPT-5.5, GPT-5.6 Luna and Nimble 9B on 2026-09-30; the others on 2026-09-21. — = no number. Iryx and Jev 1.13 were both measured with a new connection per request, but Jev’s time includes the round trip over the internet and the TLS handshake, and Iryx’s doesn’t: the gap between them isn’t only the model. In batches, each batch’s time divided by its cases counts for every case in the batch. * Includes the agent or program that called the model, and its start-up: not the time of one request. The general-purpose models got the same prompt; another setup could give another number. Nimble 9B is an open model (Apache 2.0), with 9 billion parameters; q4 is its compressed version. Its time was measured on a small home graphics card, where the full version doesn’t fit and part of it sits in CPU memory: it doesn’t represent its time on a card suited to the model. Claude is a trademark of Anthropic; GPT, of OpenAI. The names only identify the system measured. None of these companies took part in or endorses this comparison.
How to read it. Text has no neutral answer key: this one is ours, and another key could give another order. On math, the math is the truth: Iryx got all 33. On text, the general-purpose models are ahead on every test, with a single tie (GPT-5.6 Luna, on blind test 1) and the widest gap on blind test 2, with a much longer time per case, measured in batches and including program start-up: not the time of one request. With 100 questions, a gap of up to ~7 points is within the margin of error; with 192 or 225, up to ~5; on the 33 math questions, the margin is much wider. On blind test 2, Iryx is behind all of them: 73%, against 87% to 97% for the other nine. Nimble 9B, open and much smaller, is behind Iryx on the main test’s text, ties on blind test 1 and is ahead on blind test 2. All on short texts in Portuguese; the other systems may have changed since.
[ iryx ]
[ general purpose ]
A decision in milliseconds, for fractions of a cent. That changes what can be automated. Illustrative examples. The confidence values show the response format and are not measured results.
[ wearables · health ]
Fast and affordable enough to decide on every sensor reading, not once a day.
[ business · systems ]
Thousands of events per minute, triaged the moment they arrive.
[ development ]
One API call, one typed response. The software acts on the decision directly, with no text to interpret.
[ time · routine ]
Messages, tasks and requests triaged automatically. Only what needs you reaches you.
[ security ]
When every second counts, the decision comes out instantly. Below the confidence threshold, a person checks.
Billed per decision, never per token. No free plan: Try it now shows the product without an account. Pay yearly and get 2 months free. Prices in Brazilian reais. Until the platform is live, reservations are by email.
[ plus ]
R$ 110 per month
or R$ 1,100 per year, with 2 months free
[ ultra ]
R$ 490 per month
or R$ 4,900 per year, with 2 months free
[ business ]
R$ 1,990 per month
or R$ 19,900 per year, with 2 months free
[ dedicated business ]
from R$ 4,900 per month
For high volume or a custom setup. The format is defined together with the Iryx team.
Reserve↗[ pay as you go ]
R$ 0.033 per 1,000 decisions
No monthly fee. Minimum top-up of R$ 50.
Reserve↗Prices in Brazilian reais (R$). The unit is the decision: a request with three decisions counts as three. Each plan's overage costs that plan's price per 1,000. Business volume is confirmed when we talk.
[ calculator ]
Type how many decisions you expect per month. The estimate uses only the monthly prices on this page and shows the cheapest plan for that volume.
At this volume, Business dedicated (from R$ 4,900 per month) may work out better. Talk to the Iryx team.
Estimate at monthly prices; pay yearly and get 2 months free. Each plan's overage costs that plan's price per 1,000. Pay as you go: R$ 50 minimum top-up. Business volume is confirmed when we talk.
A direction, not a measured result.
[ brand consistency ]
Iryx checking whether a piece follows its design system (color, type, spacing), with a confidence number. The piece’s data goes in, as the design file describes it; out comes yes or no, and which rule it broke.
[ art direction at scale ]
Small, fast decisions that keep one aesthetic across thousands of pieces.
[ second opinion ]
Software agents checking with Iryx before they act, with the confidence band saying when to stop and ask a person.
[ brief triage ]
The written brief goes in; out come the type of piece, the format and the urgency, each with its confidence. On the low band, a person checks before production starts.
[ art direction from the brief ]
The brief and the brand’s art directions go in; out comes the most likely direction, with its confidence, and whether the piece needs the art director. On the low band, the art director decides.
[ model evolution ]
We keep improving Iryx to get more text right and to be more accurate where its confidence drops today. No date and no promised number: a new version only goes live after it’s measured.
How Iryx was made:
The dashboard isn't live yet. For now, everything starts by email.
Tell us what you want to decide, and at what volume.
Plus, Ultra, Business or pay as you go. The Iryx team helps you estimate your monthly decision volume.
With your key, one POST /v1/decide request carries the context and the decisions. Back come the value, the confidence and the spectrum of each one.
A decision model. It takes a context and closed questions (pick one option, yes or no, a level on a scale) and returns each answer with measured confidence and the full spectrum. It doesn't write text.
EYE-riks. Spelled with a Y: I, R, Y, X.
Each question answered. A request with three decisions counts as three. Billing is per decision, never per token.
No. Try it now, on this page, shows the product with no account and no credit. For ongoing use, plans start at R$ 110 per month, and pay as you go starts with a R$ 50 top-up.
On text. On the main test without math, it scores 79.2% (192 questions); Jev 1.13 and the general-purpose models, 84.4% to 88.5% (the open, smaller Nimble 9B, 76.6%). On blind test 2, it scores 73%, against 87% to 97% for the other nine systems (100 questions). On blind test 1, the general-purpose models score 86% to 96%, against 86% (100 questions): GPT-5.6 Luna ties. Where there’s math, it got all 33. When confidence lands in the low band, check before acting. Measured on 2026-09-22, 2026-09-23 and 2026-09-30, on our own answer key.
Today, the API doesn't store the context or the answers: it reads, decides in memory and replies. The log for each request keeps only operational data: the time, the method, the route, the status, the request id, the number of decisions and the duration, plus the error type when something fails. The hosted service isn't live yet; its rules will be in the terms of use before the first customer.
The API takes any text. The numbers on this page were measured on short texts in Portuguese; in English, quality hasn't been measured yet.
Classified.