# OneHot

> OneHot is an API from Sqwish Labs for small decision models called kettei
> (fat, chub, core and dot, largest first). `POST /v1/decide` takes some context and one
> or more questions, each with a fixed set of allowed answers, and returns a probability for
> every answer. You can fine-tune a family of models on a dataset that grows, and call the
> family by its name or one of its models by its number.

The OpenAPI file at [/openapi.json](/openapi.json) is the source of truth, so do not invent
fields. Every error has a stable code and a message that says what to change. The
reference is at [docs.sqwish.ai](https://docs.sqwish.ai/api-reference/overview). Paths below
are relative to your server's address, which the examples call `$SQWISH_BASE_URL`. The hosted
API is at `https://api.sqwish.ai`.

## Models

The model ids are `kettei-fat`, `kettei-chub`, `kettei-core` and
`kettei-dot`, largest first. `GET /v1/models` lists them all, plus
the fine-tunes you have trained; `GET /v1/families` lists your families. A size you can call now has `"status": "ready"`. One whose
engines are all away has `"status": "unavailable"` and `"availability_reason": "no_engine"`,
and a request for it gets a retryable `engine_unavailable` or another size's answer.
Fine-tunes train on core and chub; fat and dot can't be fine-tuned yet. The local preview runs
core only.

## Auth

Send `Authorization: Bearer onehot_sk_...` with every `/v1` request. Create and revoke keys in the
console; a new key is shown once. `GET /v1/account/keys` lists them. A key reads its account but
can't change it: keys, billing and settings change in the console, and only the console downloads
a copy of the account's data. Local protected routes also require a key.

Some routes need no key. `POST /v1/playground/decide` takes the same body as
`/v1/decide`, runs base models only, allows up to 8 decisions, is rate-limited and never
stores anything. It is for the console, behind Cloudflare's browser check, so code calls
`/v1/decide` instead. `POST /v1/lint` checks how questions are asked, without running a
model. `GET /v1/recipes` lists the ready-made requests, without their labelled examples.

## Run a decision

Each decision has an `id`, a `question` and its allowed outcomes. `kind` is `single` (one
named outcome, the default), `binary` (yes or no) or `ordinal` (levels `"0"`, `"1"` and so
on, with one `rubric` line per level). Ask related questions about the same context in one
request: they are read together, each knowing the others, and the context is billed once.
Send a question in a request of its own for an answer that doesn't depend on the others.
Keep outcomes and JSON keys in their original order, because order changes what the model
reads.

```sh
curl -s "$SQWISH_BASE_URL/v1/decide" \
  -H "Authorization: Bearer $SQWISH_API_KEY" -H "Content-Type: application/json" \
  -d @- <<'EOF'
{
  "model": "kettei-core",
  "context": "Our billing export has failed for two days. We need it for today's audit.",
  "decisions": [
    {"id": "team", "kind": "single", "question": "Which team should handle this ticket?",
     "outcomes": {"billing": "Charges, invoices, and financial exports.",
                  "technical": "Other application failures.", "sales": "New purchases."}},
    {"id": "urgent", "kind": "binary",
     "question": "Does the ticket state a deadline within one day?",
     "outcomes": {"yes": "An explicit deadline is today or within the next 24 hours.",
                  "no": "The deadline is later, absent, or unclear."}},
    {"id": "impact", "kind": "ordinal",
     "question": "How severe is the reported impact on the customer's work?",
     "outcomes": ["0", "1", "2"],
     "rubric": ["Minor inconvenience; the task can still be completed.",
                "The task is impaired, but a workable alternative is stated.",
                "An important task is blocked and no workable alternative is stated."]}
  ]
}
EOF
```

core answered this request on 26 September 2026. An excerpt, rounded to four places:

```json
{
  "id": "dec_963b7180cbd24e69a09870f8",
  "model": "kettei-core",
  "decisions": {
    "team": {"kind": "single", "top": "billing", "p_top": 0.9268,
             "probabilities": {"billing": 0.9268, "technical": 0.0539, "sales": 0.0192}},
    "urgent": {"kind": "binary", "top": "yes", "p_top": 0.9124,
               "probabilities": {"yes": 0.9124, "no": 0.0876}},
    "impact": {"kind": "ordinal", "top": "2", "expected_level": 1.7931,
               "probabilities": {"0": 0.0507, "1": 0.1056, "2": 0.8437}}
  }
}
```

Every decision returns `kind`, `probabilities`, `top` and `p_top`, and ordinal ones add
`expected_level`. A decision's optional `policy` picks an action after the model runs:
`"policy": {"action_costs": {"approve": {"yes": 1200, "no": 0}, "review": {"yes": 15, "no": 15}},
"abstain_below": 0.6}` gives each action a cost for every outcome and abstains below the
threshold. The answer then adds `"policy": {"action": ..., "expected_costs": {...}}`; without
a policy there is no action. `probabilities` always stays the model's own distribution, and
probabilities are conditional on the outcomes you declared. The base models' calibration is
not measured yet, so do not read `p_top` as an accuracy; a model you train reports its
calibration error (`ece`) on its held-out rows. `GET /v1/recipes` lists recipes, complete requests for common jobs, and
`GET /v1/onehot` returns them with their labelled examples.

Before you ship new questions, send the same body to `POST /v1/lint`. It returns
`{"warnings": [{"decision": ..., "rule": ..., "severity": ..., "message": ...}]}` for
problems such as a negative question or a choice with no way out, and an empty list when it
finds none.

## Families and models

The `model` field accepts:

| `model` | What answers |
|---|---|
| `kettei-core` | the base model |
| `router` | the model your family `router` serves, while the family is switched on (409 `family_off` or `family_empty` otherwise, 410 `family_deleted` once you deleted it) |
| `router@3` | model 3, while its own switch is on (409 `model_off` otherwise) |

A family is a line of models for one task, each trained from the one before on one dataset
that grows in versions. Only a model that passed its gate gets a number. An adapter id, or
anything after `@` but a number, such as `router@latest`, is never called: 422
`invalid_reference`.
Switch with `POST /v1/families/router/switch {"on": true}` or
`POST /v1/models/router@3/switch {"on": true}`. Each answering model takes one of your tier's
places; past them a switch gets 409 `live_limit`, and `switch_off` frees places in the same step.
A switch stays on until you switch it off, however long nobody uses it, so switch off what you
no longer call to free its place. Families and models show `used_at`: the latest production call
with an API key by that name or number, evaluation naming it, or switch-on. Calls to a family
don't move its models' own `used_at`. `switches_off_at` is null, as nothing switches off for
going unused, and `off_reason` is `owner` with `off_at` once you switch it off.
`places` in `GET /v1/account` counts the models that answer and those kept warm, each once.

A family that is switched on serves its latest model, the highest number that isn't deleted or discarded.
When a model joins, the server loads and probes it at its next check, every 30 seconds, then
moves the family to it. `GET /v1/families/router` shows `served`, the model the family
serves, and `latest`, its newest model, with its `models` and `jobs`.
A move takes a place only when the model the family served keeps answering by its own switch,
or when the family served nothing before. With no place free the family keeps its model and
moves at a later check. `GET /v1/families/router` shows that as `move`, such as
`{"to": "router@5", "waiting_for": "place", "reason": "live_limit"}`. A kept-warm family can wait
for `warm_place`, when the account keeps as many models warm as its tier allows, or
`warm_capacity`, when the model's size has no kept-warm place free. `reason` is the code of the
refusal that holds the move. `engines` means the model loads or waits to load again, and `null`
means no move waits. `moved` says when the family last moved by itself and `from` which model,
until it serves another. The owner gets one email when a family moves by itself, and one
when a move first waits, saying what holds it. A model that fails to load or to answer its probe
leaves the family where it was. After an outage that passes, the next check tries the move
again; after any other failure, it waits longer each time, up to an hour, or until an engine
joins or restarts. A family that is switched off doesn't move. Switching it on serves its latest
model or, if that fails to load, another model the family served; with none, it stays off. To
stop a family answering, switch it off. To take it back to an older model, roll it back. Keep
warm follows the family to its new model, and nothing is switched off to make room. A move that
would keep one more model warm than your tier or the model's size allows waits, as a move waits
for a place.

When a family answers, the response says which model ran:
`"model": "router", "version": "router@3"`. `GET /v1/families/router` lists every model of
the line with its held-out results and gate.

## When a model can't answer

If the model you asked for can't answer right now, for example because its adapter isn't
loaded or its engine is restarting, a base size that is ready answers instead: a fine-tune's
own size first, then the others smartest first (fat, chub, core, then dot). The response then adds `"fallback": {"from": "router", "to": "kettei-core", "reason": "adapter_not_loaded"}`.
A base model hasn't learned your labels, so check for `fallback` where that matters. Send
`"fallback": "none"` to get the error instead, for example when you evaluate one model, or
a list of up to four families or models, such as `router@2`, to try in order, passing over
any that doesn't answer calls: a family that is off, serves no model or was deleted, a model
that is off, or a deleted or discarded model. A wrong request never falls back.

## Grow a family

Read the prices first: `pricing` in `GET /v1/onehot`. A fine-tune is priced per million
training tokens (`fine_tune_million_tokens_usd`): the tokens it trains on, each time it reads them,
so two passes count a dataset twice, with a minimum per job (`fine_tune_minimum_usd`). The job
shows its count as `split.trained_tokens` and holds the fee while it runs, and is charged only if it
succeeds. A dataset over `dataset_free_rows` examples costs
`dataset_day_usd` a day from when it is made, or grows past that size, until you delete it,
whether or not you train on it. `DELETE /v1/datasets/{id}` deletes it for good; it can't be
undone. A refusal for credit is `402 insufficient_credit`. Before you start a fine-tune, send
the same body without `max_price_usd` to its estimate: `POST /v1/fine-tuning/jobs/estimate`, or
`POST /v1/families/{family}/continue/estimate` for a family's next model. It returns the exact
`price_usd` the job would hold now, its `trained_tokens`, `steps`, `seed`, `passes` and
`minimum_applied`, and queues and holds nothing. Send that price as `max_price_usd`, and a job
that would cost more, as a price, the dataset or the family changed, is refused with
`409 price_changed`.
[Laboratory costs](https://docs.sqwish.ai/guides/accounts-and-usage#laboratory-costs)
explains each fee.

1. Create a dataset. `POST /v1/datasets {"name": "...", "examples": [...]}` takes JSON up to
   1 MiB. `POST /v1/datasets/upload` takes JSONL up to 16 MiB with
   `Content-Type: application/x-ndjson`. Each example is a decide request without `model`,
   plus `targets` (a probability for every outcome, summing to 1) or `rewards` (a reward for
   every outcome). Use at least two different contexts. Read the stored rows back with
   `GET /v1/datasets/{dataset_id}/rows`. Add rows later as the dataset's next version with
   `POST /v1/datasets/{dataset_id}/versions {"examples": [...]}`, or upload JSONL
   (`Content-Type: application/x-ndjson`) or CSV (`text/csv`) to
   `POST /v1/datasets/{dataset_id}/versions/upload`. Added rows ask the dataset's own
   decisions, and each context keeps its side of the split.
2. Start a family: `POST /v1/fine-tuning/jobs {"dataset_id": "ds_...", "family": "router"}`.
   The name is new: lowercase letters and digits with single hyphens, starting with a letter.
   `from` is core (`kettei-core`), chub (`kettei-chub`) or one of your models
   such as `"router@2"`, and core when left out. fat and dot can't be fine-tuned yet: 503
   `model_unavailable`.
   Leave out `hyperparameters.steps` for two passes over the rows (20 to 2,000 steps). A family starts
   on 100 to 10,000 training rows (422 `too_few_rows` or `too_many_rows`) and needs 30
   held-out decisions its gate can judge, or 409 `not_enough_evaluation_data`; started from one
   of your models, those about contexts that model's line trained on don't count. An account
   can have four fine-tunes queued, running or cancelling (`effective_limits.fine_tunes` of
   `GET /v1/account` says how many yours can): another gets 429 `tenant_queue_full`,
   and a full training queue 429 `queue_full`, each with `Retry-After`.
3. Poll `GET /v1/fine-tuning/jobs/{job_id}` until `status` is `succeeded`, `failed` or
   `cancelled`. A run takes minutes to hours.
4. Read `metrics.gate`. The new model is compared with the model it trained from, or its base
   when the family started from one, on the same held-out decisions, split into earlier rows (older data) and new
   rows, so a big new batch cannot hide forgetting. When `passed` is false, `reasons` says why,
   and the model joins nothing: the job's `outcome` is `gate_failed`. A model that passes joins
   the family with the next number, such as `router@1`.
5. Add rows to the dataset, then train the family's next model with
   `POST /v1/families/router/continue {}`. It trains from the family's latest model on the new
   rows, with as many earlier rows again mixed in to reduce forgetting. Send `"latest"`, the id
   of the latest model you saw, and a family that moved on is refused with 409 `family_changed`.
   It is refused with 409 `no_new_data` without new rows, and `not_enough_evaluation_data`
   while they hold fewer than 30 held-out decisions. `GET /v1/families/router` says first, in
   `continue`: its `state` (`new_data`, `retry` after a run that failed or was cancelled, or
   the reason it can't), the dataset `version`, `added_at` and `added_from` (the family whose
   page added it, or null), the `rows` it would take and
   `price_usd`: 0 when it would be free, otherwise null, as the fee depends on the body you
   send; its estimate gives the exact fee. Each of the family's
   `jobs` shows its `outcome`, `price_usd`, `charge` and, for one that made no model, `reason`.

Send an `Idempotency-Key` header when you create a dataset or a job, and send the same key if
you retry. `GET /v1/fine-tuning/jobs/{job_id}/events` says why a model failed its gate.

## Evaluate models on a dataset

POST /v1/evaluations {"dataset_id":"ds_...","models":["kettei-core","router@2"]}
pins the selected models and queues only missing model/data pairs. Each model is one
a call could name: a base model, a family that is on, or a model whose switch is on. Existing
successful or active evaluations are reused; the response gives queued and reused counts.
Normal datasets reserve a fixed 20% of context groups for every model evaluation and
fine-tune. Training seeds do not change this set. Create a dataset with
"evaluation_fraction":1 (or upload JSONL with X-Evaluation-Fraction:1) to evaluate all rows
and prohibit fine-tuning and prompt tuning. Dataset purpose is immutable.

GET /v1/evaluations?dataset_id=ds_... lists results and progress.
GET /v1/evaluations/{id} reads one job. POST /v1/evaluations/{id}/cancel stops it;
POST /v1/evaluations/{id}/retry resumes failed or cancelled work from saved predictions.
New predictions use normal inference credits with fallback disabled. Reuse is free.
Fine-tuning jobs expose dataset_evaluation for the output model; verified training
predictions are reused when they cover the dataset's exact evaluation rows.

## Label real requests with teacher models

Teacher models can label requests you have no labels for: upload them as JSONL to
`POST /v1/labeling/sources`, start `POST /v1/labeling/jobs`, review
`GET /v1/labeling/jobs/{job_id}/results`, then send the rows you accept to
`POST /v1/labeling/jobs/{job_id}/finalize`, which makes a dataset for step 2 above.
`GET /v1/labeling/teachers` says whether this server has teachers set up.

## Tune the wording instead

With 50 to 1,000 training rows and no appetite for training, tune the words, not the weights.
`POST /v1/prompt-tuning/jobs {"dataset_id": "ds_...", "decision": "team"}` rewrites that
decision's question, outcome descriptions and rubric, and tests the result once on rows it
never saw. Outcome ids never change. Poll `GET /v1/prompt-tuning/jobs/{job_id}`. When
`result.verdict` is `improved`, use `result.tuned` as the decision; otherwise keep yours.
`result.label_issues` lists rows whose labels contradict the wording. A job that can't prove
an improvement on its first test of those rows costs nothing; a rerun on the same rows is
stricter and charged. `model` is a base model, or a fine-tuned model while it answers calls:
a family that is on, or a model by its number whose switch is on.

## Roll back, delete and keep warm

- `DELETE /v1/models/router@2` deletes a model for good: it answers and runs nothing
  (410 `model_deleted`), stops counting against storage, and its family's line keeps only its
  number, id, parent, status `deleted` and `deleted_at`. Name it by
  number: an id or a family's name gets 422 `invalid_reference`. A family's latest model
  can't be deleted (409 `latest_model`): roll the family back instead. A model its family
  serves, or whose own switch is on, can't be deleted (409 `model_answers`): switch
  it off if it is on; a model its family serves can be deleted once the family, switched on,
  has moved to its latest model.
- `POST /v1/families/router/rollback {"to": "router@2", "discard": ["router@3"]}` takes a
  family back to an older model and discards every newer one for good (410
  `model_discarded`). `discard` names exactly those models, or 409 `rollback_changed`. A
  family that is on loads `to` first. Its fine-tunes still training are cancelled, and so are
  evaluations still running of the models it discards.
- `DELETE /v1/families/router` deletes a family. One no model ever joined frees its name, and
  a family named so again starts a clean line. Otherwise every model of its line goes too, and
  the name is never given again (409 `family_exists`): calls, reads and changes of router get
  410 `family_deleted`. Its dataset stays, and its jobs stay listed with their charges and
  `family_deleted: true`. 409 `family_answers`, listing them, while router or one of its models
  answers calls: switch them off first. 409 `family_training` while it trains, and 409
  `version_in_use` while an evaluation scores one of its models. Send
  `{"dataset_id": "ds_...", "latest": "ft-..."}` as you saw them, or 409 `family_changed`.
- Deleting a family's dataset leaves its models answering, but the family can't continue
  (410 `dataset_deleted`, and `continue.state` `dataset_deleted`): start a new family from its
  latest model.
- `POST /v1/families/router/keep-warm {"enabled": true}` keeps the model the family serves
  loaded between requests, and follows the family when it moves. Where models load on
  demand it is billed by the hour at its base model's price (`pricing.keep_warm_hour_usd` in
  `GET /v1/onehot`), per started minute.
- `POST /v1/models/router@2/keep-warm {"enabled": true}` keeps that model loaded on its own,
  whichever model the family serves. Its own switch must be on (409 `model_off`), and switching
  it off stops its keep warm. A model kept warm both ways takes one of the tier's `kept_warm`
  places and is billed once.
- A model that isn't kept warm and hasn't been used for a while is unloaded to make room for
  others, and loads again on its next call, usually within a second. That decision waits up to
  10 seconds for it to load, then gets a base size's answer marked
  `fallback` with reason `model_warming`, or, with `"fallback": "none"`, a 503
  `model_warming` with `Retry-After: 5`. The load carries on either way. That request charged
  nothing and its idempotency key keeps this answer, so send it again after 5 seconds with a
  new key. Keep warm keeps a model loaded, subject to capacity, so its calls don't wait for a
  load; there is no fixed latency guarantee.

## Moving from TypeSafe's Jev

An existing Jev request works unchanged at `POST /v1/systemone`. Point it at your OneHot
server and send a OneHot key. Jev's model ids (`jev-latest`, `jev-1.13.0`) run on
`kettei-core`, and the answer's `model` field says so. A field OneHot doesn't model
is rejected with a 422 that names it, rather than ignored.

```sh
curl -s "$SQWISH_BASE_URL/v1/systemone" -H "Authorization: Bearer $SQWISH_API_KEY" \
  -H "Content-Type: application/json" -d '{"model": "kettei-core",
  "state": "My parcel arrived broken.",
  "questions": {"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"}}}'
```

New code should use the native shape, which says the same thing in OneHot's words:

```sh
curl -s "$SQWISH_BASE_URL/v1/decide" -H "Authorization: Bearer $SQWISH_API_KEY" \
  -H "Content-Type: application/json" -d '{"model": "kettei-core",
  "context": "My parcel arrived broken.",
  "decisions": [{"id": "refund", "kind": "binary", "question": "Is the customer asking for a refund?"}]}'
```

`state` becomes `context` and `questions` become `decisions`. `noul`, `choice` and `score`
become `binary`, `single` and `ordinal`, and `criteria` become outcome descriptions or an
ordinal `rubric`. `/v1/systemone` keeps Jev's `confidence` so thresholds tuned on Jev keep
working. `/v1/decide` returns every probability with `p_top`, and adds no confidence score
of its own.

`POST /v1/systemone/translate` does the conversion for you. Send it a Jev request and it
returns the native one, with notes on anything that changed, including how a Jev confidence
threshold becomes a `p_top` threshold. It runs no model and needs no key.

## Errors

Every error has the same shape:
`{"error": {"code": "...", "message": "...", "retryable": false, "request_id": "...", "details": [...]}}`.
`retryable` says whether the same request can succeed later. Retry only when it is true,
after the `Retry-After` header when there is one, and with the same `Idempotency-Key`. A 503
can go either way: an engine blip is retryable, a size this server doesn't serve is not. On
a 422, `details` lists each field that failed with its path, so fix those fields rather than
sending the same body again. If a switch or keep-warm request fails without a
response, read the family or model before you repeat it.

## Tools for coding agents

- An MCP server with tools to run and check decisions, tune their wording, train and keep
  models warm. Servers installed with the `mcp` extra serve it at `/mcp` (streamable HTTP,
  same API key as a bearer token; a browser session gets 401 `api_key_required` there), and
  `onehot mcp` runs it over stdio. Setup for Claude Code, Codex and Cursor is in
  `docs/agents.md` in the workbench repository.
- Python SDK: `pip install sqwishai`, then `from sqwishai import OneHot` and
  `OneHot(url, api_key=...)`. Its source is `sdks/python` in the workbench repository.
- TypeScript SDK: `npm install sqwishai`, then `import { OneHot } from "sqwishai"` and
  `new OneHot(url, { apiKey })`. It has no dependencies. Its source is
  `sdks/typescript/onehot.ts` in the workbench repository.
- Agent skill in `skills/onehot/SKILL.md`.
- `POST /v1/hooks/claude-code` lets a kettei model check Claude Code's tool calls before
  they run. Setup is in `docs/claude-code-hook.md`.

## Reference

Dataset, job and labelling lists come back newest first with `has_more` and `next_cursor`;
send that value as `cursor` to get the next page. Dataset rows and labelling results page by
`offset` and `limit` instead.

- [OpenAPI schema](/openapi.json)
- [API reference](https://docs.sqwish.ai/api)
