← A11y Desk / API
Tokens

Drive A11y Desk from your own code

Everything the web page does is available over HTTP: post one component's markup — HTML, JSX or TSX, a Vue or Svelte template — and get back either a WCAG 2.2 AA audit of it (findings tied to success criteria, a keyboard map, the checks only a human can make, and a verdict) or an accessible rewrite of the same component (the whole thing again, fixed, with a change log and the things markup alone cannot carry). The natural use is a CI job that audits the components a pull request touched and fails the build when a blocker appears, or a batch that walks a component library overnight and produces the backlog nobody has had time to write down.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{ "ok": true,  "data":  { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "details": { ... } } }

Send your token as Authorization: Bearer … on every call. The app slug travels in the body of /guest as {"slug": "a11y-desk"}; after that the token itself carries the app, so a run needs only Authorization, Content-Type: application/json and the Idempotency-Key described in step 5.

The request body for /estimate, /run and /run-stream is the input object itself — not wrapped in anything. A body of {"input": {…}} returns 200 and quietly hides every field from the model, so a run that looks fine comes back describing nothing. Guard it the way the app's own client does: refuse to send anything that is not a plain JSON object.

Error codes

codestatuswhat to do
unauthorized401The token is missing, malformed or expired. Get a new one from the token page.
payment_required402The balance is below min_credits, or a guest token tried a metered run. Call /estimate first, then sign in and top up.
validation_error400The body is not a plain JSON object, or a field is the wrong type — code must be a string. task, code and framework are required; a missing one comes back as a warning from /estimate.
not_found404Unknown job id, or the app slug does not exist. Check the id you polled with.
rate_limited429Too many requests. Back off and retry; do not tight-loop a poll.
internal500A server-side failure. Retry with the same Idempotency-Key so you are not billed twice.

Replaying an Idempotency-Key with a changed body is not a retry and is rejected rather than billed. When you change the input — a different component, a tightened constraints line — bump the attempt suffix on the key instead.

Structured fields travel as JSON strings. The platform checks every body against a declared list of scalar fields, so prescan_facts and findings_to_fix are declared as strings: the web page sends JSON.stringify(prescan_facts) and JSON.stringify(findings_to_fix). The samples below show them as plain objects for readability; that still works and the model reads both forms, but /estimate then returns a should be string warning for each. Encode them to keep the warnings list empty, so a real mistake such as an {"input": {…}} wrapper stands out.

The field to get right first: task

A11y Desk is one app with two lanes, and task is what chooses between them. It is the first field of every request body:

taskwhat comes backextra input fields
"audit"The component inspected against WCAG 2.2 AA: findings tied to a success criterion each, passes, manual_checks a person still has to run, a keyboard_map, a score of severity counts, and a verdict of blocked, needs_work or ready_for_manual_testing.none
"rewrite"The same component again, fixed: rewritten_code, a changes log with verbatim before and after slices, unresolved items the markup cannot settle, behaviour_notes for what needs script rather than markup, and a test_script to check it by hand.findings_to_fix, constraints

The system prompt routes on task and never blends the two contracts in one reply. A missing or unrecognised task is not an error: the model answers the closest lane, names the lane it actually answered in the reply's lane field, and says so in notes_on_input. So branch on lane in the reply, never on the task you believe you sent.

The two lanes chain, and the handoff is the point of the app. Run audit, take the findings you intend to fix, and send them back as the findings_to_fix of a rewrite run over the same code. Each entry is a small object — {"id": "F-001", "sc": "1.3.1", "summary": "…"} — and the rewrite then reports, in changes[].fixes, which of those ids each edit closed. The web page does this with one button, and a second button runs the rewritten markup back through the audit lane, which is the only honest way to see whether the rewrite actually moved anything.

1. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool — it reads the same storage the app itself uses and prints the token for you.

A guest token can call /me and /estimate. Auditing a component and rewriting one are both metered, so either lane needs a personal token from signing in.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://a11y-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="aut_YOUR_TOKEN"
#
# To mint a guest token from the command line instead. A guest token is enough for
# /me and /estimate; an audit or a rewrite run needs a personal token.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" \
  -d '{"slug": "a11y-desk"}'
# {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}

2. A tiny client

One helper that adds the headers, unwraps data and raises on error.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="a11y-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://a11y-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

3. Check the session and the balance

GET /me tells you whether the token is a guest or a person, and what the balance is. The object is small and carries exactly three things: subject_type — guest or user — subject_id, and credits, the wallet balance. There is no username in it, so "signed in" is subject_type === "user" and nothing else; a guest can price a run but cannot start one. Compare credits against min_credits from the next step before you run, so a shortfall surfaces as your own clear message rather than a 402.

call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}

# A guest token answers the same call with subject_type "guest" and cannot run
# either lane. Gate on it before you spend a poll loop finding out:
call me | grep -q '"subject_type":"user"' || {
  echo "sign in at https://a11y-desk.skillsafe.ai/tokens.html first" >&2
  exit 1
}

4. Price the run — free

The input object is exactly what the app's own form submits. It is always a JSON object — never a bare string, never wrapped in an input key:

fieldtypemeaning
taskstring, required"audit" or "rewrite". The lane. It routes the prompt, and a missing or unknown value degrades to the closest lane rather than failing — the reply names what it answered in lane and in notes_on_input.
codestring, requiredOne component's markup, pasted as text: HTML, JSX or TSX, a Vue or Svelte template. This is the run's only evidence — every finding quotes a snippet out of it, and every rewrite is this text again. The browser clips to 40,000 characters from the middle, keeping the beginning and the end, and leaves a marker in place of the cut. Clip the same way if you send more and keep the marker: the prompt keys on it, refuses to claim anything about the missing middle, and mentions the cut in notes_on_input. Send one component, not a whole page — a component is what both lanes are scoped to.
frameworkenumhtml, jsx, vue, svelte, angular or unknown. Detected client-side from the markup and overridable by the user. It matters more than it looks: it decides whether the fix is for= or htmlFor=, whether a handler is onclick or a framework binding, and what the rewrite is allowed to emit.
context_hintstring, optionalUp to 600 characters about where this component lives and what it has to support — "Login form on the marketing site; Tailwind; evergreen browsers only". It is what lets the audit tell a real constraint from a preference.
prescan_factsobject, always presentWhat the app's free local scanner found before the run. Shape and honesty note below. Always send the object, even when issues is empty.
rewrite lane only
findings_to_fixobject[], optional[{"id": "F-001", "sc": "1.3.1", "summary": "…"}] — the findings from an earlier audit run over the same code. This is the handoff: every id you send should come back in some changes[].fixes or be explained in unresolved. Omit it and the rewrite fixes what it finds itself, and every fixes array comes back empty.
constraintsstring, optionalUp to 600 characters of what the rewrite may not do — "keep the class names; no new dependencies". Use it for the house rules that would otherwise make a correct rewrite unmergeable.

prescan_facts, and why a script should still send it

In the browser this object is computed for free, before the run and without spending a credit, by the app's own static scanner — it never leaves the page, and the model sees only the summary below:

{
  "framework": "jsx",
  "stats": { "lines": 41, "elements": 23, "interactive": 6, "images": 2,
             "form_fields": 3, "headings": 2, "landmarks": 1, "components": 2 },
  "headings": ["h1 Dashboard", "h4 Recent activity"],
  "components": ["Button", "Icon"],
  "issues": [
    { "id": "S-001", "rule": "img-alt", "sc": "1.1.1", "level": "A", "severity": "serious",
      "line": 12, "snippet": "<img src=\"/icon.svg\">", "message": "img has no alt attribute" }
  ],
  "manual_hints": ["forms", "dialog", "icons", "colour_in_css"],
  "clipped": { "cut": 0 }
}

stats is a shape count of the component. headings is the heading outline as the scanner read it, and components names the capitalised elements whose internals are not visible in the paste — an audit cannot claim anything about what <Button> renders, and saying so is what keeps it honest. manual_hints lists what static review cannot settle in this particular markup, and clipped.cut is how many characters the middle clip removed.

Each entry in issues is one mechanical hit, with severity in blocker, serious, moderate or minor, sc as the dotted success-criterion number and level as A, AA or AAA. The scanner's rule ids are a fixed vocabulary, one WCAG 2.2 criterion each:

img-alt  img-alt-redundant  input-image-alt  svg-name  control-name  control-name-dynamic
link-no-href  link-text-generic  label-missing  placeholder-as-label  label-empty
div-click  role-button-no-tabindex  role-button-no-keys  tabindex-positive
aria-hidden-focusable  aria-label-no-role  aria-role-invalid  aria-attr-invalid
aria-ref-missing  id-duplicate  heading-skip  heading-empty  heading-multiple-h1
html-lang  title-missing  iframe-title  table-headers  fieldset-legend
radio-group-no-fieldset  autoplay  viewport-zoom  contrast  focus-outline-none
target-size  dialog-name  dialog-modal  role-heading-level  meta-refresh
nested-interactive  li-outside-list  motion-no-reduce  marquee

An API caller does not have to reproduce any of that — there is no scanner to call over HTTP, and the model reads code either way. What a script should do is send the object in the shape shown, with the arrays empty: {"framework": "html", "stats": {…}, "headings": [], "components": [], "issues": [], "manual_hints": [], "clipped": {"cut": 0}}. The field is always present in a real request, the prompt is written against its presence, and an empty issues array is a legitimate, meaningful value: it says the caller ran no scan, not that the component is clean.

Be honest with yourself about what an empty scan gives up. The contract that makes the facts worth sending is this: every id in prescan_facts.issues comes back exactly once in coverage, with a status of confirmed, downgraded, dismissed or merged. A mechanical scanner is allowed to be wrong, and dismissed with a reason is the honest answer for a false positive — a different thing from silence. Send no issues and coverage comes back empty, so you lose the reconciliation, not the audit. If you have your own linter output, map it into this shape and send it: it is the strongest check in the whole reply.

/estimate creates no job and charges nothing. It returns the model binding — model is gpt-5.6-terra, model_alias is gpt-terra, and markup_bps — plus sponsor_enabled and the reservation: hold_credits is what gets held, and min_credits is the balance you must clear to start at all. The hold is a reservation, not the price. It prices the full output cap, so the charged_credits on the settled job is usually far lower. Budget against hold_credits, report against charged_credits.

Price each lane separately. A rewrite body carries the whole component again and prices differently from an audit of the same paste, and a prescan_facts carrying forty issues is forty issues' worth of input tokens.

# One small, thoroughly inaccessible login form is the worked example throughout.
# Lane A - the audit. prescan_facts carries what a local scan found; both ids come
# back in `coverage`.
INPUT='{"task": "audit", "code": "<form>\n  <div>Email</div>\n  <input type=\"email\" placeholder=\"Email\">\n  <div class=\"btn\">Log in</div>\n</form>", "framework": "html", "context_hint": "Login form on the marketing site; Tailwind; evergreen browsers only", "prescan_facts": {"framework": "html", "stats": {"lines": 5, "elements": 4, "interactive": 1, "images": 0, "form_fields": 1, "headings": 0, "landmarks": 0, "components": 0}, "headings": [], "components": [], "issues": [{"id": "S-001", "rule": "label-missing", "sc": "1.3.1", "level": "A", "severity": "serious", "line": 3, "snippet": "<input type=\"email\" placeholder=\"Email\">", "message": "input has no associated label"}, {"id": "S-002", "rule": "placeholder-as-label", "sc": "3.3.2", "level": "A", "severity": "moderate", "line": 3, "snippet": "<input type=\"email\" placeholder=\"Email\">", "message": "placeholder is used in place of a label"}], "manual_hints": ["forms", "colour_in_css"], "clipped": {"cut": 0}}}'

call estimate "$INPUT"
# {"ok":true,"data":{"model":"gpt-5.6-terra","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":2652,"min_credits":310,"sponsor_enabled":false}}
#
# estimate is FREE. It creates no job and charges nothing. hold_credits is what
# gets RESERVED; charged_credits on the settled job is normally much lower.

# Lane B - the rewrite of the SAME component, carrying two findings the audit
# produced. This is the handoff the web page's button performs.
FIX_INPUT='{"task": "rewrite", "code": "<form>\n  <div>Email</div>\n  <input type=\"email\" placeholder=\"Email\">\n  <div class=\"btn\">Log in</div>\n</form>", "framework": "html", "context_hint": "Login form on the marketing site; Tailwind; evergreen browsers only", "constraints": "keep the class names; no new dependencies", "findings_to_fix": [{"id": "F-001", "sc": "1.3.1", "summary": "Email input has no associated label"}, {"id": "F-002", "sc": "2.1.1", "summary": "A div is the submit control: not focusable and not operable by keyboard"}], "prescan_facts": {"framework": "html", "stats": {"lines": 5, "elements": 4, "interactive": 1, "images": 0, "form_fields": 1, "headings": 0, "landmarks": 0, "components": 0}, "headings": [], "components": [], "issues": [{"id": "S-001", "rule": "label-missing", "sc": "1.3.1", "level": "A", "severity": "serious", "line": 3, "snippet": "<input type=\"email\" placeholder=\"Email\">", "message": "input has no associated label"}, {"id": "S-002", "rule": "placeholder-as-label", "sc": "3.3.2", "level": "A", "severity": "moderate", "line": 3, "snippet": "<input type=\"email\" placeholder=\"Email\">", "message": "placeholder is used in place of a label"}], "manual_hints": ["forms", "colour_in_css"], "clipped": {"cut": 0}}}'

call estimate "$FIX_INPUT"   # price each lane separately

5. Run it, then poll

POST /run returns a job_id; poll GET jobs/{job_id} until status is succeeded or failed. The result JSON is the string at data.output.output. The terminal job also carries charged_credits — the real price — and the truncated flag.

Always send an Idempotency-Key. The web app builds it as a11y-desk:<task>:<input hash>:a<attempt> and so should you. Three parts, three reasons:

A retried request carrying the same key returns the same job instead of billing a second run, which is what makes a CI retry safe after a network blip. It also means the audit-then-rewrite handoff is cheap to re-run: the audit half replays, and only the rewrite is new work.

# The key is slug:task:hash:attempt. A retried request with the same key returns
# the SAME job instead of billing a second run.
KEY="a11y-desk:audit:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

# Poll until the job reaches a terminal status.
while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# The terminal job looks like this:
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"audit\",\"title\":\"Login form - WCAG 2.2 audit\", ...}"},
#   "charged_credits":588,"truncated":false}}

# The rewrite half of the handoff is the same call with the other body and its
# own key - never reuse the audit's key for it.
FIX_KEY="a11y-desk:rewrite:$(printf '%s' "$FIX_INPUT" | shasum -a 256 | cut -c1-16):a1"

6. Or stream it

POST /run-stream is the same call over server-sent events, and it takes the same Idempotency-Key. Each delta event carries {"text": "..."}, a chunk of the result JSON, and the final done event carries status, charged_credits — the real price, normally a fraction of the hold — and the truncated flag.

Read the SSE yourself. What a client receives depends on where it is: a command-line reader like the ones below gets real delta events, while the same endpoint sends a page in a browser tick heartbeats instead — so a JavaScript callback wired to deltas never fires there, and any progress display, streaming preview or partial-recovery path built on it is dead code in a browser. Parse the event stream in your own reader, as the samples here do, and treat a run with no deltas at all as normal rather than as a stall: wait for done, or fall back to /run and polling.

The practical tip: do not try to parse the partial JSON to drive a progress display — watch for key names arriving in the accumulating text instead. In the audit lane the appearance of "findings", then "passes", then "manual_checks", then "keyboard_map", then "score" is what advances the stage from reading the markup to naming what a person still has to test. In the rewrite lane the sequence is "rewritten_code", "changes", "unresolved", "behaviour_notes", "test_script" — and because rewritten_code arrives first and is by far the largest string, a rewrite looks stalled for most of its run unless you say so. Substring matching on the quoted key name is enough, and it costs nothing.

# Server-sent events. From the command line each `delta` carries a chunk of the
# JSON; the final `done` event carries the status, charged_credits and truncated.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: delta
# data: {"text":"{\"lane\":\"audit\",\"title\":\"Login form"}
# event: delta
# data: {"text":" - WCAG 2.2 audit\",\"findings\":[{"}
# event: done
# data: {"status":"succeeded","charged_credits":588,"truncated":false}
#
# A browser is sent `tick` heartbeats instead of deltas. Read the stream in your
# own reader, or fall back to /run and polling.

7. Parse the result

data.output.output is a string holding one JSON object — unwrap twice. The web app strips an optional code fence, takes everything from the first { to the last }, parses that, and only then reads fields. Doing the same two things — the fence strip and the outer-brace slice — is what makes a caller robust against the small variations a model produces around an otherwise clean object.

Here is an abbreviated audit reply for the login form above, structurally complete:

{
  "lane": "audit",
  "title": "Login form - WCAG 2.2 AA audit",
  "summary": "The form has two controls and neither is usable as written. The email field has no programmatic label, so a screen reader announces an unnamed edit box, and the submit control is a div, which no keyboard user can reach or activate. Fixing the submit control is the difference between a form that cannot be completed and one that merely reads badly.",
  "component": "login form",
  "verdict": "blocked",
  "verdict_reason": "The submit control is a plain div: it is not focusable and not operable by keyboard, so the form cannot be submitted at all without a mouse.",
  "findings": [
    { "id": "F-001", "sc": "1.3.1", "sc_name": "Info and Relationships", "level": "A",
      "severity": "serious", "principle": "perceivable", "affects": ["screen_reader"],
      "line": 3, "snippet": "<input type=\"email\" placeholder=\"Email\">",
      "problem": "The text 'Email' sits in a sibling div and in the placeholder, neither of which is a programmatic label. A screen reader user hears 'edit, blank' and has no way to know what the field wants; the placeholder also disappears the moment they start typing.",
      "fix": "Give the input an id and associate a real label with it, and drop the placeholder or make it an example value rather than the field's name.",
      "fix_code": "<label for=\"email\">Email</label>\n<input id=\"email\" type=\"email\" name=\"email\" autocomplete=\"email\" required>" },
    { "id": "F-002", "sc": "2.1.1", "sc_name": "Keyboard", "level": "A",
      "severity": "blocker", "principle": "operable",
      "affects": ["keyboard", "screen_reader", "motor"],
      "line": 4, "snippet": "<div class=\"btn\">Log in</div>",
      "problem": "The submit control is a div. It is not in the tab order, it exposes no button role, and Enter and Space do nothing on it, so a keyboard user, a screen reader user and anyone using switch access cannot log in at all.",
      "fix": "Use a real button element and keep the class so the styling is unchanged.",
      "fix_code": "<button type=\"submit\" class=\"btn\">Log in</button>" }
  ],
  "passes": [
    { "sc": "1.1.1", "sc_name": "Non-text Content", "note": "the component carries no images or icons, so nothing here needs a text alternative" }
  ],
  "manual_checks": [
    { "id": "M-001", "sc": "2.4.7", "what": "focus is visible on every control",
      "how": "Tab through the form; each control must show a focus indicator with at least 3:1 contrast against its background. The Tailwind class on the button may remove the default outline." },
    { "id": "M-002", "sc": "1.4.3", "what": "the label and the button text meet contrast",
      "how": "Sample the rendered colours with a contrast checker: 4.5:1 for the label text, 3:1 for the button's border against the page." }
  ],
  "keyboard_map": [
    { "element": "Email field", "keys": "Tab", "expected": "focus lands in the field and its name is announced", "status": "missing" },
    { "element": "Log in control", "keys": "Tab, Enter, Space", "expected": "submits the form", "status": "missing" }
  ],
  "score": { "blocker": 1, "serious": 1, "moderate": 0, "minor": 0 },
  "coverage": [
    { "id": "S-001", "status": "confirmed", "ref": "F-001", "note": "" },
    { "id": "S-002", "status": "merged", "ref": "F-001", "note": "The placeholder-as-label hit is the same defect as the missing label, so it is one finding, not two." }
  ],
  "credential_seen": false,
  "notes_on_input": ""
}

And the rewrite reply for the same component, carrying the two findings the audit produced:

{
  "lane": "rewrite",
  "title": "Login form - accessible rewrite",
  "summary": "Both findings are closed in markup alone: the email field gets a real label and an autocomplete token, and the div becomes a submit button with its class kept. One thing is left open, because the component has nowhere to put a sign-in error.",
  "framework": "html",
  "rewritten_code": "<form>\n  <label for=\"email\">Email</label>\n  <input id=\"email\" type=\"email\" name=\"email\" autocomplete=\"email\" required>\n  <!-- TODO: render the sign-in error here and move focus to it -->\n  <button type=\"submit\" class=\"btn\">Log in</button>\n</form>",
  "changes": [
    { "id": "C-001", "sc": "1.3.1",
      "what": "Replaced the text div with a real label bound to the input by id, and moved the field name out of the placeholder.",
      "before": "<div>Email</div>",
      "after": "<label for=\"email\">Email</label>",
      "fixes": ["F-001"] },
    { "id": "C-002", "sc": "2.1.1",
      "what": "Turned the styled div into a submit button, keeping the btn class so nothing changes visually.",
      "before": "<div class=\"btn\">Log in</div>",
      "after": "<button type=\"submit\" class=\"btn\">Log in</button>",
      "fixes": ["F-002"] }
  ],
  "unresolved": [
    { "id": "U-001", "sc": "3.3.1",
      "why": "The component has no error region, and the markup alone does not say where a failed sign-in message should appear or what it says.",
      "placeholder": "<!-- TODO: render the sign-in error here and move focus to it -->" }
  ],
  "behaviour_notes": [
    "A submit button posts the form; if the sign-in is handled in script, the handler must listen for submit on the form rather than for click on the button, or keyboard submission is lost again.",
    "The error region added at the TODO needs aria-live=\"polite\" or focus moved into it, otherwise a screen reader user never learns the attempt failed."
  ],
  "test_script": [
    { "step": "Tab once from the top of the form", "expect": "focus lands in the email field and 'Email, edit' is announced" },
    { "step": "Tab again, then press Enter", "expect": "focus is on the Log in button and the form submits" },
    { "step": "Submit with an empty field", "expect": "the browser's own required-field message names the Email field" }
  ],
  "coverage": [
    { "id": "S-001", "status": "confirmed", "ref": "C-001", "note": "" },
    { "id": "S-002", "status": "merged", "ref": "C-001", "note": "The placeholder went away with the same change." }
  ],
  "credential_seen": false,
  "notes_on_input": ""
}

The verdict rule

The audit verdict is not free-form, and the page re-derives it rather than trusting it — treating the findings list as authoritative and warning when the returned verdict disagrees:

blocked                    if any finding has severity "blocker"
needs_work                 else if any finding is "serious" or "moderate"
ready_for_manual_testing   otherwise - minor findings only, or none at all

Do the same. A gate that reads verdict alone can be talked out of failing by a reply that lists a blocker and then calls itself needs_work. And note what the top verdict means: blocked is not "bad", it is "a user of one assistive technology cannot complete this component's purpose at all" — an unlabeled field they cannot identify, a control unreachable by keyboard, a modal with no name and no way out. ready_for_manual_testing is the ceiling, not a pass: it says static review found nothing more, and the manual_checks list is now the work.

While you are there, check score against findings. It is four counts by severity over the same list, so it is the cheapest possible test that the reply is internally consistent, and a reply whose score disagrees with its own findings is one whose verdict you should not trust either.

The checks the page runs on a rewrite

A rewrite is easier to get subtly wrong than an audit, because it produces one long string that looks plausible whatever is in it. These are the five checks the web app performs on every rewrite reply, all of them cheap and all of them worth copying:

The last one needs a scanner you do not have over HTTP — so the API equivalent is to re-audit: post rewritten_code back into the audit lane and compare score with the first run's. That costs a second run and is worth it in CI, where the alternative is merging a rewrite nobody checked.

# The result JSON is a string inside the envelope, so unwrap it twice.
RESULT=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])')

# Strip an optional code fence and keep the outer {...}, then re-derive the verdict.
printf '%s' "$RESULT" | python3 - <<'PY'
import json, sys
raw = sys.stdin.read().strip()
if raw.startswith("```"):
    raw = raw.split("\n", 1)[1].rsplit("```", 1)[0]
obj = json.loads(raw[raw.index("{"):raw.rindex("}") + 1])

sev = [f.get("severity") for f in obj.get("findings", [])]
want = ("blocked" if "blocker" in sev
        else "needs_work" if ("serious" in sev or "moderate" in sev)
        else "ready_for_manual_testing")
print(obj["lane"], obj.get("verdict"), "re-derived:", want)
if obj.get("verdict") != want:
    print("WARNING: the reply's verdict disagrees with its own findings", file=sys.stderr)
PY

The output contract

One JSON object, no prose and no code fence around it. First the envelope both lanes share:

keytypemeaning
laneenumaudit or rewrite — the lane the model actually answered, and therefore which contract the rest of the object follows. Branch on this, not on the task you sent.
titlestringA short name for this run, taken from the component's own subject — "Login form - WCAG 2.2 AA audit".
summarystringTwo to four sentences: what the component is and the one thing the reader must know about it.
coverageobject[]{id, status, ref, note} — the reconciliation table, one entry per prescan_facts.issues[] id, exactly once. status is confirmed (it became a finding or a change, and ref is that F- or C- id), downgraded (real, but less severe than the scanner claimed — ref set, note says why), dismissed (a scanner false positive — ref is "" and note says why) or merged (folded into another finding, whose id is ref).
credential_seenbooleantrue when the pasted markup looked like it carried a password, API key or token — a hard-coded bearer in a fetch call, a key in a data attribute. The reply then repeats no part of the value anywhere. Treat it as a signal to rotate, and keep it out of your logs.
notes_on_inputstring"" when there is nothing to say. Carries: that the middle of the markup was clipped, that the task was missing or unrecognised and which lane was answered instead, and that the detected framework disagrees with the one you sent.

The audit body

keytypemeaning
componentstringWhat the markup is, in the component's own words — "login form", "pricing table", "nav drawer".
verdictenumblocked, needs_work or ready_for_manual_testing. Derived, not chosen — see the verdict rule above, and re-derive it rather than trusting it.
verdict_reasonstringOne sentence naming the finding or the fact that decided the verdict.
findingsobject[]{id, sc, sc_name, level, severity, principle, affects, line, snippet, problem, fix, fix_code}. Ids run F-001 upward. sc is the dotted success-criterion number and sc_name its title; level is A, AA or AAA; affects names who is blocked; line is an integer or null; snippet is verbatim from your code, up to 200 characters; problem and fix are one to three sentences each; fix_code is the corrected markup and may be "" when the fix is not a markup change.
passesobject[]{sc, sc_name, note} — criteria this component actually meets. Shipping what passed is what makes the findings list auditable rather than a wall of complaints.
manual_checksobject[]{id, sc, what, how}, ids M-001 upward. The work static review cannot do: contrast against rendered colours, a visible focus ring, a screen-reader pass. how is the actual procedure, not a restatement of the criterion.
keyboard_mapobject[]{element, keys, expected, status} with status in ok, missing, unknown. One row per interactive element: what keys should do, and whether the markup as written delivers it. unknown is the honest answer when the behaviour lives in script you did not paste.
scoreobject{blocker, serious, moderate, minor} — counts over findings, and they must match it. The cheapest internal-consistency test in the reply.

The rewrite body

keytypemeaning
frameworkenumSame enum as the input. The rewrite keeps the framework you pasted; anything else has to be explained in notes_on_input.
rewritten_codestringThe whole component again, fixed, as one JSON string. Not a diff and not an excerpt — it is meant to replace the file's component wholesale.
changesobject[]{id, sc, what, before, after, fixes}, ids C-001 upward. before is verbatim from your input and after verbatim from rewritten_code, each up to 200 characters, so a reviewer can find both ends of every edit. fixes lists the findings_to_fix ids the change closes, and is [] when you sent no findings.
unresolvedobject[]{id, sc, why, placeholder}, ids U-001 upward. What the markup alone cannot settle — the meaning of a chart, the text of an error nobody wrote. placeholder is the marker left in rewritten_code, and it appears there verbatim.
behaviour_notesstring[]What needs script rather than markup: a focus trap, an Escape handler, a live region, focus returned to the control that opened a dialog. These are the things a rewrite cannot do for you and must not pretend to have done.
test_scriptobject[]{step, expect} — a short by-hand pass over the rewritten component, in the order a person would actually perform it.

The enums

Ids are zero-padded to three digits and sequential from 001, with the letter naming what they are: S- a scanner issue you sent, F- an audit finding, M- a manual check, C- a rewrite change, U- something the rewrite could not resolve. That is what makes the handoff mechanical — a findings_to_fix entry is an F- id, and the changes[].fixes that comes back names it.

8. Use it in CI

The worked example: a job reads a component out of the repository, audits it, prints every finding with its success criterion, and exits non-zero when the component is unusable — when the verdict is blocked, or when any finding carries a severity of blocker or serious. Both halves of that test matter. The verdict alone misses a reply that lists three serious findings and calls itself needs_work; the severities alone miss nothing, but re-deriving the verdict from them and comparing is what catches a reply that is internally inconsistent.

Let moderate and minor through and report them. A gate that fails on every minor finding in a living component library becomes noise people learn to skip, and the findings that actually block someone are then invisible. Derive the Idempotency-Key from the component's contents so a re-run of the same commit replays the same job instead of re-billing, and bump the attempt suffix only when the markup really changed.

The script below sends prescan_facts with empty arrays, which is the honest thing for a caller with no local scanner — and it means coverage comes back empty, so the gate leans on the findings and the verdict rather than on reconciliation. If your project already runs an accessibility linter, map its output into the issues shape and send it: the reconciliation check is the strongest signal in the reply, and it is how you find out which of your linter's hits are false positives.

The workflow that runs it, as GitHub Actions YAML:

name: accessibility

on:
  pull_request:
    paths:
      - "src/components/**"

jobs:
  a11y:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - name: Audit the login form
        env:
          SKILLSAFE_TOKEN: ${{ secrets.SKILLSAFE_TOKEN }}
        run: python3 ci/a11y-gate.py src/components/LoginForm.html
#!/bin/sh
# a11y-gate.sh - fail the build when a component cannot be used.
# Usage: a11y-gate.sh src/components/LoginForm.html
set -eu

BASE="https://api.skillsafe.ai/v1/app-api"
TOKEN="$SKILLSAFE_TOKEN"
FILE="$1"

# Build the body with python3 so the markup is JSON-escaped correctly.
BODY=$(python3 - "$FILE" <<'PY'
import json, sys
markup = open(sys.argv[1], encoding="utf-8").read()
print(json.dumps({
    "task": "audit",
    "code": markup,
    "framework": "html",
    "context_hint": "Component from the application repository, audited in CI",
    "prescan_facts": {
        "framework": "html",
        "stats": {"lines": markup.count("\n") + 1, "elements": 0, "interactive": 0,
                  "images": 0, "form_fields": 0, "headings": 0, "landmarks": 0,
                  "components": 0},
        "headings": [], "components": [], "issues": [],
        "manual_hints": [], "clipped": {"cut": 0},
    },
}))
PY
)

KEY="a11y-desk:audit:$(printf '%s' "$BODY" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$BODY" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  STATE=$(curl -sS "$BASE/jobs/$JOB" -H "Authorization: Bearer $TOKEN")
  S=$(printf '%s' "$STATE" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$S" = "succeeded" ] && break
  [ "$S" = "failed" ] && echo "$STATE" && exit 2
  sleep 2
done

printf '%s' "$STATE" | python3 - <<'PY'
import json, sys
job = json.load(sys.stdin)["data"]
raw = job["output"]["output"].strip()
if raw.startswith("```"):
    raw = raw.split("\n", 1)[1].rsplit("```", 1)[0]
obj = json.loads(raw[raw.index("{"):raw.rindex("}") + 1])

for f in obj.get("findings", []):
    print(f"{f['severity']:8} {f['sc']:6} {f['id']}  {f['problem'][:90]}")

sev = [f.get("severity") for f in obj.get("findings", [])]
hard = [s for s in sev if s in ("blocker", "serious")]
print(f"verdict={obj.get('verdict')} charged={job.get('charged_credits')} truncated={job.get('truncated')}")
if job.get("truncated"):
    sys.exit("the reply was truncated - top up and re-run, do not gate on a prefix")
sys.exit(1 if obj.get("verdict") == "blocked" or hard else 0)
PY

Truncation and partial results

When the balance sits between min_credits and hold_credits, the run is not refused: it executes with a reduced output cap and comes back with truncated: true on the finished job and on the streaming done event. What you hold then is a prefix of the reply, not the reply. In the audit lane the first findings may be complete while manual_checks, keyboard_map, score and coverage are missing or cut mid-string; in the rewrite lane rewritten_code is written first and is the longest string in the object, so what gets lost is the change log, the unresolved list and the behaviour notes — everything that tells you what the rewrite did.

Check the flag before you treat a reply as complete. The shape of the object will not tell you: a truncated audit that still carries two findings parses cleanly and looks like a small audit, and its score and verdict — the two fields a gate reads — are exactly the fields most likely to be missing. A truncated rewrite is worse, because rewritten_code cut off mid-element is still a string, and a person skimming a diff can merge it.

The right response is a retry, not a repair: top up, or resubmit a smaller component — one control group rather than a whole form — with the attempt suffix on the Idempotency-Key incremented so the new body is not a replay of the old key. Repairing truncated JSON by appending closing braces produces something that parses and is not what the model meant; in this app it produces an audit whose score disagrees with the findings above it, or a rewrite whose change log describes edits its code does not contain.

One more honest limit: code is clipped from the middle at 40,000 characters, and the reply says so in notes_on_input. An audit of a clipped component is an audit of its beginning and its end, and a rewrite of one is not safe to merge at all — the middle it never saw would come back missing. Both lanes are scoped to one component for this reason. For a whole page, split it at its own component boundaries, run each part, and let the manual_checks lists tell you what still has to be tested across the assembled page: focus order, landmark structure and the reading order between components are exactly the things no per-component run can settle.