IRYX™company
Meet:
Iryx 1.0
Say:
EYE-riks, with a Y
Product:
Iryx API v1

Iryx 1.0 decides.

Our first decision model. You send a situation and the questions you want answered, with the possible options. Iryx picks the answer to each question and tells you how sure it is. It doesn't write text: it decides.

Measured on Iryx's own answer key on 2026-09-23, on short texts in Portuguese. Next to nine other systems: see where it does well and where it falls behind.

Response:~74 msmedian per request, through the API
Math:99.4%exact, 970 of 976
Main test:82%225 questions
Blind tests:73–86%2 tests, new domains

How it works

Three steps. You ask, Iryx decides and tells you how sure it is, and you choose when the system can act on its own.

  1. [ 01 · you send ]

    The context and the questions.

    The context is the situation: a message, a record, a JSON. With it go the questions, each with its possible options. There are three kinds: pick one option, yes or no, a level on a scale.

    > context I bought sneakers from your store last week and they came in the wrong size. [...] Can I exchange them by Friday? > questions
    subject?pick one: exchange or return · late delivery · payment · cancellation · product question
    needs a reply today?yes or no
    customer's tone?scale: calm · worried · annoyed · furious
  2. [ 02 · iryx decides ]

    One answer per question.

    For each question, Iryx returns the chosen option, the confidence (from 0 to 1) and the spectrum: the chance of every option, not only the chosen one.

    < decisions
    subjectexchange or return · 0.93
    reply todayyes · 0.71
    toneworried · 0.95
  3. [ 03 · you set the threshold ]

    Above the threshold, it acts. Below, a person checks.

    You choose the minimum confidence for the system to act on its own. Anything below goes to a person. The band (high, medium, low) is the cut we recommend.

    example threshold: 0.80
    exchange or return · 0.93acts alone
    worried · 0.95acts alone
    reply today: yes · 0.71person checks

Example: the sneaker exchange from Try it now, a real test run on 2026-09-24 with a made-up message in Portuguese. The three kinds of question and the values are the ones from that test. The 0.80 threshold is only an example: you set it.

Glossary

Decision
Each question answered. It is also the billing unit.
Context
The situation you send: a message, a record, a JSON.
Confidence
A number from 0 to 1 that says how sure Iryx is of the chosen option.
Spectrum
The chance of every possible option, not only the chosen one.
Band
High, medium or low: the confidence cut we recommend. On low, check before you act.
Typed answer
The answer always comes in a fixed format (one of the options, yes or no, a level on the scale), which software uses directly, with no text to interpret.

Try it now

A context goes in, decisions come out. Each one with the value, the confidence and the full spectrum. The band tells you when to check.

Pick a message:

Try it now · session 001REC
See the API request POST /v1/decide

< decisions

Input:
context + decisions
Output:
value, confidence, spectrum
Case:
real test
Demo
Ready-made answers: this page doesn't call the model.
Real test
The sneaker exchange was run on 2026-09-24, with a made-up message in Portuguese. The test kept the chosen value of each decision; the rest of the spectrum shows as a sum.
Examples
The double charge and the store review are illustrative examples.
Band
High, medium or low: the confidence cut we recommend, computed on this page. On low, check before you act.

[ benchmarks · own answer key ]

Iryx 1.0 in numbers

Nearly the same accuracy as the largest models, in a fraction of the time per case. On math, it gets every one right. Ten systems, the same questions, our own answer key.

Systems:
10
Questions:
up to 425
Answer key:
own
Answers:
2026-09-21–22, 30
Measured:
2026-09-22–23, 30

Nearly the same accuracy,
a fraction of the time.

On the full main test, 82.2%, against 84.9% to 88.4% for the general-purpose models. Per case, ~90 ms; the general-purpose models, 802 to 4,747 ms, measured in batches: not the time of one request. The open Nimble 9B, 568 to 1,790 ms on a small home graphics card. The same cases for all, in the same order in every chart.

Accuracy on our own answer key. * Time includes the agent or command line that called the model, in batches: not the time of one request.

~74 msper request of ~3 decisions, median, through the API, not over the internet, on a kept-open connection; with a new connection per request, ~90 ms. The others were called other ways: see the table.
82.2%on the full main test (225 questions: math + text). The general-purpose models score 84.9% to 88.4%; Jev 1.13, 80.9%; Nimble 9B, 71.6% (68.4% for q4).
970 of 976on math, across two sealed tests, where the math is the truth. On the 33 that all ten systems took, Iryx got every one; the other nine got 14 to 31.
CH 02 · 10 seconds at median time · 10 systemsREC
00.0 sper request * in batches

How to read it

Each square is one complete case, with ~3 decisions, at each system’s median time. It illustrates time per case; it isn’t a throughput measure. Iryx: 200 requests, 2026-09-23. GPT-5.5, GPT-5.6 Luna and Nimble 9B: 2026-09-30. The others: 300 cases, 2026-09-21.

Iryx, Jev 1.13 and Nimble 9B

Per request, through the API, with a new connection per request. Jev goes over the internet, with TLS; Iryx doesn’t: the gap between them isn’t only the model. On a kept-open connection, Iryx takes ~74 ms. Nimble 9B runs locally, through Ollama, on a small home graphics card: it doesn’t represent a suitable card.

* In batches

The six general-purpose models were called in batches, by an agent or from the command line. The time is the batch time divided by its cases, start-up included: it isn’t the time of one request. In a batch, the cases come out together, at the end.

All the numbers

Accuracy of the ten systems on our own answer key, on each test
SystemMath
33
MainBlind 1
100
Blind 2
100
Text only
192
Full
225
Iryx 1.0100%79.2%82.2%86%73%
Jev 1.1357.6%84.9%80.9%87%92%
Nimble 9B42.4%76.6%71.6%86%88%
Nimble 9B q442.4%72.9%68.4%86%87%
Claude Opus 593.9%87.5%88.4%96%97%
Claude Fable 5.187.9%88.5%88.4%96%93%
GPT-6 Astra90.9%86.5%87.1%95%96%
GPT-5.6 Sol87.9%87.5%87.6%96%95%
GPT-5.6 Luna87.9%84.4%84.9%86%91%
GPT-5.584.8%87.0%86.7%92%93%

Accuracy on our own answer key; the number under each test is the question count. The full main test adds the 33 math questions to the 192 text ones. Answers from 2026-09-21 (math, main and blind 1) and 2026-09-22 (blind 2), measured on 2026-09-22 and 2026-09-23; for GPT-5.5, GPT-5.6 Luna and the two Nimble 9B, and for GPT-5.6 Sol on blind test 2, answered and measured on 2026-09-30.

Time per case and how each system was called
SystemTime per case
median
CasesHow it was called
Iryx 1.090 ms200API, not over the internet, one request at a time, new connection per request (on a kept-open connection: 74 ms)
Jev 1.13412 ms300Vendor API, over the internet, new connection per request
Nimble 9B1,790 ms—local Ollama API (/v1/systemone), one request at a time, on a small home graphics card the model doesn’t fit on whole (part in CPU memory)
Nimble 9B q4568 ms—same, in the compressed version (q4), which fits whole on the same card
Claude Opus 5802 ms*300
in 6 batches
agent, batches of 50; batch time ÷ 50
Claude Fable 5.11,515 ms*300
in 6 batches
agent, batches of 50; batch time ÷ 50
GPT-6 Astra4,069 ms*300
in 30 batches
OpenAI command line, batches of 10, low reasoning; process time ÷ 10
GPT-5.6 Sol4,747 ms*300
in 30 batches
OpenAI command line, batches of 10, low reasoning; process time ÷ 10
GPT-5.6 Luna2,714 ms*—command line (Codex), batches of 10; batch time ÷ 10
GPT-5.52,491 ms*—command line (Codex), batches of 10; batch time ÷ 10

A case is one request with ~3 decisions. Iryx measured on 2026-09-23; GPT-5.5, GPT-5.6 Luna and Nimble 9B on 2026-09-30; the others on 2026-09-21. — = no number. Iryx and Jev 1.13 were both measured with a new connection per request, but Jev’s time includes the round trip over the internet and the TLS handshake, and Iryx’s doesn’t: the gap between them isn’t only the model. In batches, each batch’s time divided by its cases counts for every case in the batch. * Includes the agent or program that called the model, and its start-up: not the time of one request. The general-purpose models got the same prompt; another setup could give another number. Nimble 9B is an open model (Apache 2.0), with 9 billion parameters; q4 is its compressed version. Its time was measured on a small home graphics card, where the full version doesn’t fit and part of it sits in CPU memory: it doesn’t represent its time on a card suited to the model. Claude is a trademark of Anthropic; GPT, of OpenAI. The names only identify the system measured. None of these companies took part in or endorses this comparison.

How to read it. Text has no neutral answer key: this one is ours, and another key could give another order. On math, the math is the truth: Iryx got all 33. On text, the general-purpose models are ahead on every test, with a single tie (GPT-5.6 Luna, on blind test 1) and the widest gap on blind test 2, with a much longer time per case, measured in batches and including program start-up: not the time of one request. With 100 questions, a gap of up to ~7 points is within the margin of error; with 192 or 225, up to ~5; on the 33 math questions, the margin is much wider. On blind test 2, Iryx is behind all of them: 73%, against 87% to 97% for the other nine. Nimble 9B, open and much smaller, is behind Iryx on the main test’s text, ties on blind test 1 and is ahead on blind test 2. All on short texts in Portuguese; the other systems may have changed since.

Iryx 1.0 or a general-purpose model?

[ iryx ]

When Iryx fits better

  • High volume: many decisions a day, billed one by one.
  • Real time: the answer has to come back in milliseconds.
  • Closed answers: pick one option, yes or no, a level on a scale.
  • Predictable cost: priced per decision, not per token.
  • Knowing when to check: the confidence tells you when a person should look.

[ general purpose ]

When a general-purpose model fits better

  • Free text: writing, summarizing, translating, explaining.
  • Conversations and open-ended tasks, with no list of options.
  • Top accuracy on text: when it matters more than time and cost. On the text tests on this page, they score higher.

Where Iryx 1.0 can work

A decision in milliseconds, for fractions of a cent. That changes what can be automated. Illustrative examples. The confidence values show the response format and are not measured results.

[ wearables · health ]

Health in real time.

Fast and affordable enough to decide on every sensor reading, not once a day.

Context:
reading out of range
Decision:
alert someone? yes
Confidence:
0.93
Action:
notify a family member or professional

[ business · systems ]

Systems that organize themselves.

Thousands of events per minute, triaged the moment they arrive.

Context:
new record in the system
Decision:
priority: high
Confidence:
0.89
Action:
route to finance

[ development ]

For people who build systems.

One API call, one typed response. The software acts on the decision directly, with no text to interpret.

Context:
application data
Decision:
option B
Confidence:
0.95
Action:
follow flow B

[ time · routine ]

More free time.

Messages, tasks and requests triaged automatically. Only what needs you reaches you.

Context:
new message
Decision:
needs you? no
Confidence:
0.91
Action:
archive and summarize

[ security ]

Response in milliseconds.

When every second counts, the decision comes out instantly. Below the confidence threshold, a person checks.

Context:
sensor event described in text
Decision:
trigger the service? yes
Confidence:
0.88
Action:
trigger the security service

Pricing

Billed per decision, never per token. No free plan: Try it now shows the product without an account. Pay yearly and get 2 months free. Prices in Brazilian reais. Until the platform is live, reservations are by email.

[ plus ]

Plus

R$ 110 per month

or R$ 1,100 per year, with 2 months free

Decisions per month:
3,600,000
Overage:
R$ 0.030 per 1,000
Reserve↗

[ ultra ]

Ultra

R$ 490 per month

or R$ 4,900 per year, with 2 months free

Decisions per month:
17,000,000
Overage:
R$ 0.029 per 1,000
Reserve↗

[ business ]

Business

R$ 1,990 per month

or R$ 19,900 per year, with 2 months free

Decisions per month:
70,000,000
Overage:
R$ 0.028 per 1,000
Reserve↗

[ dedicated business ]

Dedicated Business

from R$ 4,900 per month

For high volume or a custom setup. The format is defined together with the Iryx team.

Reserve↗

[ pay as you go ]

Pay as you go

R$ 0.033 per 1,000 decisions

No monthly fee. Minimum top-up of R$ 50.

Reserve↗

Prices in Brazilian reais (R$). The unit is the decision: a request with three decisions counts as three. Each plan's overage costs that plan's price per 1,000. Business volume is confirmed when we talk.

[ calculator ]

What does your volume cost?

Type how many decisions you expect per month. The estimate uses only the monthly prices on this page and shows the cheapest plan for that volume.

A request with three decisions counts as three.

    Estimate at monthly prices; pay yearly and get 2 months free. Each plan's overage costs that plan's price per 1,000. Pay as you go: R$ 50 minimum top-up. Business volume is confirmed when we talk.

    Where we're going

    A direction, not a measured result.

    [ brand consistency ]

    Does the piece follow the system?

    Iryx checking whether a piece follows its design system (color, type, spacing), with a confidence number. The piece’s data goes in, as the design file describes it; out comes yes or no, and which rule it broke.

    Track:
    art direction
    Status:
    research

    [ art direction at scale ]

    A thousand pieces, one aesthetic.

    Small, fast decisions that keep one aesthetic across thousands of pieces.

    Track:
    design systems
    Status:
    research

    [ second opinion ]

    The agent asks. Iryx decides.

    Software agents checking with Iryx before they act, with the confidence band saying when to stop and ask a person.

    Track:
    agents
    Status:
    internal use

    [ brief triage ]

    A request came in. What piece is it?

    The written brief goes in; out come the type of piece, the format and the urgency, each with its confidence. On the low band, a person checks before production starts.

    Track:
    production
    Status:
    idea

    [ art direction from the brief ]

    Which art direction does the brief call for?

    The brief and the brand’s art directions go in; out comes the most likely direction, with its confidence, and whether the piece needs the art director. On the low band, the art director decides.

    Track:
    art direction
    Status:
    idea

    [ model evolution ]

    New versions, more accurate answers.

    We keep improving Iryx to get more text right and to be more accurate where its confidence drops today. No date and no promised number: a new version only goes live after it’s measured.

    Track:
    model
    Status:
    ongoing

    How Iryx was made:

    How to start

    The dashboard isn't live yet. For now, everything starts by email.

    1. Make a reservation

      Tell us what you want to decide, and at what volume.

    2. Pick a plan

      Plus, Ultra, Business or pay as you go. The Iryx team helps you estimate your monthly decision volume.

    3. Call the API

      With your key, one POST /v1/decide request carries the context and the decisions. Back come the value, the confidence and the spectrum of each one.

    FAQ

    What is Iryx 1.0?

    A decision model. It takes a context and closed questions (pick one option, yes or no, a level on a scale) and returns each answer with measured confidence and the full spectrum. It doesn't write text.

    How do you say it?

    EYE-riks. Spelled with a Y: I, R, Y, X.

    What counts as a decision?

    Each question answered. A request with three decisions counts as three. Billing is per decision, never per token.

    Is there a free plan?

    No. Try it now, on this page, shows the product with no account and no credit. For ongoing use, plans start at R$ 110 per month, and pay as you go starts with a R$ 50 top-up.

    Where does it still fall behind?

    On text. On the main test without math, it scores 79.2% (192 questions); Jev 1.13 and the general-purpose models, 84.4% to 88.5% (the open, smaller Nimble 9B, 76.6%). On blind test 2, it scores 73%, against 87% to 97% for the other nine systems (100 questions). On blind test 1, the general-purpose models score 86% to 96%, against 86% (100 questions): GPT-5.6 Luna ties. Where there’s math, it got all 33. When confidence lands in the low band, check before acting. Measured on 2026-09-22, 2026-09-23 and 2026-09-30, on our own answer key.

    What happens to what I send?

    Today, the API doesn't store the context or the answers: it reads, decides in memory and replies. The log for each request keeps only operational data: the time, the method, the route, the status, the request id, the number of decisions and the duration, plus the error type when something fails. The hosted service isn't live yet; its rules will be in the terms of use before the first customer.

    Does it work in English?

    The API takes any text. The numbers on this page were measured on short texts in Portuguese; in English, quality hasn't been measured yet.

    How was it made?

    Classified.