Image-to-Accessible-HTML parsing service. Iris converts a sequential set of image files (e.g. the rendered pages of a PDF) into a single content-only, WCAG 2.2 AA accessible HTML document, using specialized per-content-type agents, a self-extending builder, and an iterative reader/copy-editor review loop.
The pipeline as implemented today runs in three phases:
- Extraction — for each page image, the
pageagent (agents/page.md) converts the whole page to an accessible HTML fragment in one vision call. The output is then verified, and corrected if the verifier objects. If the page agent names a content type a specialist would handle better, that specialist is dispatched and its output merged. Pages are independent, so they are extracted in parallel — up todefaults.extraction_concurrencyat a time (default 5, clamped to 1..16). Fragments keep submitted document order regardless of which page finishes first; lower it if your provider rate-limits you, or set1for fully serial. Across sessions,defaults.max_concurrent_runs(default 2, clamped to 1..32) bounds how many runs execute at once; further uploads wait instatus: "queued"rather than being rejected. - Assembly — fragments are joined in page order into a minimal accessible document shell
(
<html lang>,<title>,<main>) and validated with axe-core. - Review — the Reader reads the document in chunks as two views (HTML + a flattened
screen-reader view) and flags reading-order / semantic / accessibility issues, attributing
each to the source page(s) it appears on; the Copy Editor proposes fixes against just those
pages' source images; fixes are applied and the document re-linted. Loops up to
max_review_iterations(default 3).
When Iris meets content a specialist agent would handle better than the general pass, it drafts
that agent and automatically files a GitHub issue titled New agent suggestion: <type> (with
the agent code + context) on the upstream repo. Maintainers triage those issues; merged agents
become part of the shared agents/ library. (This replaces the PRD's fork+PR-on-close flow —
see Implementation notes.)
Those issues are identified by their title prefix, not by a label, and deliberately so: GitHub silently drops labels set by anyone without push access to the repo, which is most of the people this is built for. A label would therefore have been missing on exactly the issues that most needed it, with nothing to say so — and the duplicate check that filtered on it would have refiled the same suggestion every session, under a different person's name each time. If you want labels on these, add a repository rule keyed on the title prefix; it applies them as the repo rather than as the filer, so it works no matter who filed.
Requires Node.js 24+ (the service runs TypeScript directly via Node's built-in type
stripping and uses the built-in node:sqlite), and a git checkout of the agent library
(this repo's agents/ directory works). For PDF uploads, install poppler-utils
(pdftoppm/pdfinfo) — brew install poppler on macOS, apt-get install poppler-utils on
Debian/Ubuntu. (The Docker image includes it.)
git clone https://github.com/EqualifyEverything/equalify-iris
cd equalify-iris
npm install
cp .env.example .env # a model provider key; GitHub App settings are optional
cp config.example.yaml config.yaml
# load env and run
set -a; source .env; set +a
npm start # -> http://localhost:8080Or with Docker (multi-arch; Mac Mini / Linux ARM are first-class targets):
cp .env.example .env # fill in values
docker compose upCheck it's alive:
curl http://localhost:8080/v1/healthOr just open the accessible browser app at the root for a no-API walkthrough (sign in with GitHub → upload page images → convert → view the accessible HTML):
http://localhost:8080/
Deployment is configured in config.yaml (PRD §10.3). ${ENV_VAR} references are expanded
from the environment at startup; changes require a restart.
- Storage (§10.2): local filesystem + a single SQLite file by default.
agents/is a git checkout modified only bygit pullfrom upstream. - Model providers (§10.3): each agent declares a capability (
vision,structured_output,text); the deployment maps capabilities to a provider + concrete model. v1 ships OpenRouter and Amazon Bedrock adapters. Adding a provider is a small adapter implementing theModelProviderinterface insrc/providers/types.ts. Models are set per provider (default_model+per_capability), and can be overridden per agent viaproviders.per_agent— either a string (provider only) or{ provider, model }. Resolution falls back: per-agent model → providerper_capability→ providerdefault_model. Each provider also takesmax_tokens(default 32000), the per-call output ceiling. A response that stops at the ceiling is a failed call, not a short one: it arrives as a 200 with HTML cut mid-tag, which would otherwise be assembled into the deliverable as if it were genuine content. Both adapters reject it and the error names the knob to raise. - Concurrency (§9.4): two independent knobs under
defaults.extraction_concurrencyis within a run (pages in parallel);max_concurrent_runsis across sessions. Peak in-flight model calls is the product of the two, so the second is the one that bounds what the machine is doing — each run also holds a jsdom+axe instance. Uploads beyond the cap wait, in FIFO order, instatus: "queued"; the wait appears in the session's run log asrun_queued/run_dequeued(waited_ms). Nothing is rejected — the upload is already received and on disk, so a 429 would discard work the user has already paid for. The cap is global rather than per user because the resources it protects (memory, jsdom, the provider's rate limit) are global. - GitHub (§9.1): GitHub is the auth mechanism — a user is their GitHub account, and a token
is required on every call. By default the
service uses a bundled GitHub App via the device flow — no per-operator app setup, no
secret (the same approach the
ghCLI uses). Setgithub.client_idonly to point at your own GitHub App;client_secretis needed only if you enable the web redirect flow. No OAuth scope is requested at all — the app's one permission comes from installing it onupstream_repo— see GitHub is the only SSO layer.
There is no anonymous mode, no API key, and no second identity provider. Every request carries a user's GitHub token, and that token is what files the session's feedback back to the shared agent library — as an issue, under that user's own GitHub identity.
That is the sustainability model, not an implementation detail (PRD §12). The agents in
agents/ get better because sessions run against real documents and real corrections; a user who
could consume the service without contributing would be taking from a library nobody was refilling.
Requiring GitHub auth is how using Iris and improving it become the same act, and how each
contribution is credited to the person who produced it. If you would rather your users not
contribute, this is not the service to deploy.
Two consequences an operator should know before deploying:
1. The permission lives with the installation, not with your users. The token does exactly two
things: GET /user to identify the caller, and file issues on upstream_repo. Iris is registered as
a GitHub App, so the second one is granted once — by installing the app on upstream_repo with
issues: write — and users only authorize. Their consent screen requests no repository access
at all, because there is nothing left for it to ask for.
One limit worth knowing if your upstream_repo is private: a user's token is the intersection
of the installation's permissions and that user's own access, so installing the app does not give a
user access they did not already have. On a private upstream, filing works for users who can see the
repo and 404s for everyone else. Set github.issue_token if you need a private upstream to accept
contributions from users who are not collaborators — it files everything under one account, which
trades away the per-user attribution below. A public upstream_repo (the assumption here, since the
agent library is meant to be shared) has no such limit.
This replaced an OAuth App requesting public_repo, and the reason is worth stating plainly: there
is no OAuth scope meaning "open issues on one repository". public_repo was the narrowest one that
could file, and it grants read and write to every public repository the user can reach —
code, commit statuses, collaborators, webhooks — none of which Iris touches. Nothing pushes and
nothing opens pull requests. So the old consent screen asked for orders of magnitude more than the
service uses, and the app is the only way to fix that rather than merely document it.
What the user's token still carries is their identity. A user-to-server token acts as the user, so issues are filed under their own account and each contribution is credited to the person whose session produced it — the whole reason users authorize at all instead of the app filing as itself.
Two registration settings the service depends on, if you point client_id at your own app:
| Setting | Value | Why |
|---|---|---|
| Enable Device Flow | on | Off by default for a new app, and the device flow is the default deployment's only login path (it returns device_flow_disabled without it). |
| Expire user authorization tokens | off | With expiry on, user tokens last 8 hours and come with a refresh token. Nothing here persists or refreshes a credential, so turning expiry on means building refresh plumbing first. |
The misconfiguration this cannot catch at startup is the app not being installed on
upstream_repo — that state lives on github.com, not in config. It surfaces as a 403 or 404
during filing, logged with a hint saying so. Both statuses, because GitHub does not reveal
repositories a credential cannot see: an app that was never installed reads as 404 Not Found
rather than as a permissions error. (A misspelled upstream_repo looks identical, and the hint says
so rather than blaming the installation.) When issue_token is set, the hint names the service
PAT instead, since the installation governs only tokens issued to users.
If you are coming from an earlier build, three things changed, and two of them can stop a working deployment:
- A configured OAuth App id is now a hard startup failure. An
Ov…client_idis refused, because Iris no longer sends any OAuth scope: such an app would authenticate users and then be unable to file a single issue. Register a GitHub App (Iv…) and install it on yourupstream_repo, or leaveclient_idblank for the bundled one. upstream_repois no longer independent ofclient_id. Under the old OAuth App, thepublic_reposcope could file on any public repo, so leavingclient_idblank and repointingupstream_repoat your own agent library worked. A GitHub App'sissues: writecomes from its installation on one specific repository, and the bundled app is installed on this repo — so that same config now files nothing, for anyone. You need your own app installed on your repo (or ask us to install ours there). This combination warns at startup rather than failing, since we cannot see from config whether the bundled app was installed on your repo.github.oauth_scopeis gone. A config that still sets it — includingoauth_scope: none, which used to be a startup error — now starts fine and ignores the key. Delete it.
There is no user-facing migration: no one had authorized the OAuth App, and any existing authorization can be revoked at github.com/settings/applications.
2. github.issue_token is an override, and not a recommended one. Set it to a service-account
PAT and every issue is filed under that bot account instead of under the user who produced it. It is
off by default because it erases the attribution that is the point of the design. Use it only where
a deployment genuinely cannot file as its users — an org policy that forbids it, say.
It is never written to disk. The token arrives in the Authorization header, is used in memory
for the request and for the pipeline run it authorizes, and is gone when the run ends. There is no
github_token column in data/iris.sqlite and no token file — a stolen copy of the database is a
list of GitHub user IDs and logins, not GitHub access.
Two smaller things follow from that, both worth knowing:
- Identity lookups (
GET /user) are cached in memory for 5 minutes, keyed by the token, so a revoked token keeps working for up to that long. The cache is bounded (10,000 entries, oldest evicted) and entries are not renewed on use — deliberately, so that a busy token cannot outlive its revocation indefinitely. It is empty on restart. - Because nothing is stored, there is nothing to rotate, re-encrypt or purge when a user revokes access. Revocation at github.com is the whole mechanism.
If you have a data/iris.sqlite from an earlier build, delete it. Tokens were stored in a
github_token column once, and there is no migration — every user re-authorizes from scratch. The
service refuses to start against such a file and names the fix, rather than adopting it: the old
table's github_token TEXT NOT NULL would survive CREATE TABLE IF NOT EXISTS, so first-time
logins would fail with a SQLite constraint error returned as 401 unauthorized (users who already
had a row would keep working, which makes it look like flaky GitHub auth rather than a schema
mismatch) — and the claim above would be false for that file, since it still holds live plaintext
tokens for everyone who ever logged in. Delete it rather than archiving it; users lose only their
session history.
All endpoints are under /v1 and (except auth and health) require
Authorization: Bearer <github_token>.
| Method & path | Purpose |
|---|---|
GET /v1/health |
Liveness probe |
GET /v1/auth/github/start |
Begin OAuth (web clients) |
GET /v1/auth/github/callback |
OAuth callback → returns access token |
POST /v1/auth/github/device |
Begin device flow (CLI clients) |
POST /v1/auth/github/device/poll |
Poll device flow (send { "device_code": ... }) |
GET /v1/me |
Current GitHub user + config |
GET /v1/sessions |
List the caller's sessions |
POST /v1/sessions |
Create a session, upload images and/or PDFs (multipart/form-data) |
GET /v1/sessions/{id} |
Poll status |
GET /v1/sessions/{id}/output |
Fetch the HTML when ready |
POST /v1/sessions/{id}/feedback |
Submit feedback, trigger a re-run |
POST /v1/sessions/{id}/close |
Finalize the session and clean tmp |
GET /v1/sessions/{id}/logs |
Fetch the run log (ndjson) |
GET /v1/sessions/{id}/diagnostics |
Timing/health summary (phase + per-call durations, in-flight/hung call) |
Full copy-pasteable bash/curl walkthrough of every endpoint: docs/API.md.
To prove the endpoints work end-to-end (mock GitHub + mock model, no credentials needed):
./test/e2e.sh.
Example — create a session (order of images parts is the processing order, §9.2):
curl -X POST http://localhost:8080/v1/sessions \
-H "Authorization: Bearer $TOKEN" \
-F "images=@page-001.png" \
-F "images=@page-002.png" \
-F 'config={"max_review_iterations": 3}'Then poll GET /v1/sessions/{id} until status is ready_for_review, fetch
GET /v1/sessions/{id}/output, and POST /v1/sessions/{id}/close to finalize.
agents/ # the agent library: page.md (the general pass), feedback.md,
# and specialists dispatched by name (§7.4 v1.2)
src/
config.ts # config loader (${ENV} expansion)
providers/ # ModelProvider interface + openrouter & bedrock adapters
agents/loader.ts # loads agent .md files, pins git SHA (§7.3)
pipeline/
orchestrator.ts # runs the phases, persists results, drives learning
extraction.ts # per-page vision pass (+ verify, correct, specialist merge)
assembly.ts # joins fragments into the document shell
review.ts # reader -> copy editor -> re-lint loop (scoped image payload)
pageindex.ts # page-number index shared by the reader + feedback scoping
lint.ts # axe-core in jsdom (color-contrast disabled, see §4)
flatten.ts # screen-reader text view, used by reader + coverage
feedback.ts # verify / scope / classify / train + regression gate
memory.ts # per-agent example bank of learned corrections
regression.ts # fixture capture + pruning on close
contribute.ts # drafts suggested agents, files issues
util/queue.ts # bounded FIFO run queue (cross-session concurrency cap)
auth/ # GitHub OAuth + device flow + bearer middleware
github/ # auto-files labeled agent-suggestion issues
store/ # node:sqlite metadata store + on-disk session layout (§8.1)
routes/ # /v1 endpoints
index.ts # server entry point
data/ # sessions/, tmp/, and the SQLite DB (created at runtime)
Where v1 diverges from the PRD (read these before assuming a PRD section describes the code — tracked in #30):
-
Three phases, not five (§6). Triage and Reconciliation are not implemented, and the Builder Agent does not draft session-scoped agents into
tmp/<id>/agents/. Extraction is a single general page agent rather than triage → per-region fan-out; the fan-out was removed because it duplicated output for nested structures like forms. Reconciliation additionally cannot run until extraction emits fragment edge data (it currently emits none).Reconciliation's within-page job also no longer exists: it was there to clean up after the fan-out, and one page now yields one fragment from one agent, so there are never two fragments competing to represent the same content. Across pages the problem is real and open — a paragraph or table can span a page break, and the page agent notes the cut-off edge rather than joining anything (§7.6 v1.2).
-
One agent per page, not one per content type (§7.4 v1.2). The PRD's nine per-content-type agents (
paragraph.md,table.md,formField.md, …) have been deleted, and this is the decision on whether the agent library is the product: it is, but the library is not a taxonomy of content types. Those nine were not merely unused, they were unreachable through every path that can reach an agent file — dispatch declines each of their names before the file is looked up, onlypage.mdis ever trained, and the contribution filter blocks the same names — so no fixture, lesson or prompt improvement could ever accrue to one. Nine prompt files that cannot run are worse than none: they read as the live extraction path to anyone openingagents/.Seeing the whole page is the capability, so per-region fan-out is not coming back: nine agents re-rendering one image produced two representations of one thing (a
<form>and a<table>for the same fields) and then needed a reconciliation phase to remove a duplication the architecture had just created — at nine times the cost and latency of the single call that already produces the answer.What is left is specialization that earns its place:
page.mdas the general, trainable pass, plus specialists for content a whole-page pass demonstrably handles worse, dispatched by name and merged in.chartDataAgent.mdis the shape — reading precise values off a chart's axes into a data table is a different task, needs its own long contract, and would bloat the page prompt for every page containing no chart. Aparagraphspecialist is not that shape; "wrap prose in<p>" is one line of the page prompt. This is also why the context pressure that motivates splitting agents up is answered per-capability rather than per-content-type: a specialist's contract is loaded only for the pages that need it, whereas nine near-duplicate prompts relieve nothing.The nine type names survive as data (
STANDARDinsrc/pipeline/contribute.ts), which is what declines a suggestion the page pass already covers and what keeps it from being re-filed as a new agent to build. That list was never a mirror of the library — it is the boundary of what one whole-page call handles — so it stays data rather than a directory listing, and dropping atable.mdintoagents/does not start splicing a second table over the page's own.The names are matched case-insensitively, through one shared normalizer used by both the dispatch decline and the contribution filter. A suggestion's name is prose a model wrote, not a filename (
STANDARDitself spells one entryformField), so"Table"is ordinary output. While the nine files existed,agents/Table.mdresolved on a case-insensitive volume and absorbed it; with them gone, an exact-match filter would draft an agent and file a public issue on the upstream repo — under the user's own GitHub identity — for a type the page pass covers. -
No provenance comments in the output (§7.4/§7.7). The PRD specifies
@source/@agent/@fragmentwrappers preserved into the final HTML. Iris delivers clean content-only HTML instead: the comments leak pipeline internals into a document meant to be handed to end users, and every consumer would have to strip them. Provenance is recorded in the run log (GET /v1/sessions/{id}/logs) rather than in the deliverable.@unresolvedis emitted when the review loop hits its iteration cap with issues outstanding (§7.11). -
Contributions are issues, not PRs (§7.13/§9.2). Instead of fork+PR-on-close, when the extractor flags content a specialist would handle better, Iris drafts that agent and files a
New agent suggestion: <type>GitHub issue with the agent code + context; feedback that generalizes files anAgent update proposal: <agent>issue the same way. Simpler to triage, and it needs no write access to a fork — so nothing forks and nothing pushes. Consequently the PRD'spending_prsandprs_openedresponse fields, theskip_prsparameter and thefork_repofield on/v1/meare not part of the API. Issues are filed with the logged-in user's token, which is required, and the point;github.issue_tokenoverrides that with a service account, at the cost of the attribution. -
Review issues are attributed by page, not by
@sourceregion (§7.8/§7.9). The PRD's issue format references@sourceregion ids from the per-region fan-out, which extraction no longer produces and which are stripped from the deliverable anyway (§7.4 v1.1). Issues instead carrypages: number[]— the source pages the Reader matched the offending content to, from an index of page-number + extracted-HTML excerpt. Attribution is what scopes the Copy Editor's image payload (below); the two-view (HTML + flattened) cross-check is implemented as specified.
Places where the PRD left a decision open, and where v1 intentionally stops:
-
runs/<run-id>vssessions/<session-id>. The PRD references both (§7.3/§7.5 vs §8.1). This implementation treats the run id as the session id and writes the log,agent-updates.md, etc. undersessions/<session-id>/, matching the authoritative layout in §8.1. (Two files in that tree,new-agents.mdandprs.md, are not written at all — they belong to the withdrawn fork-and-PR flow; see §8.1 v1.2.) -
Reader chunking (§7.8). Chunks use a fixed character budget with overlap rather than a literal 30%-of-context computation, since the per-model context window is not exposed through the provider abstraction. The two-view (HTML + flattened) cross-check is implemented as specified.
-
Color-contrast lint. Output is content-only with no styling (§4), so axe-core's
color-contrastrule is disabled — it cannot be assessed without rendering and is out of scope. -
Duplicate ids are linted for three separate ways (§7.7 v1.2). Obsolete as a conformance criterion is not the same as harmless here: this document is assembled from independently extracted pages, so a duplicate id is the specific defect concatenation produces, and it breaks navigation rather than conformance. Two
id="fn-1"means everyhref="#fn-1"reaches the first one, so a footnote reference on a later page silently goes to the wrong note while the link still looks like it works. Covering that takes three rules, because axe splits the check by what the element is and each rule skips the others' elements:duplicate-id(elements nothing references and nothing focuses) andduplicate-id-active(focusable ones) are both taggedwcag2a-obsolete— WCAG 2.2 dropped 4.1.1 — so the tag filter would skip them and each is enabled by name.duplicate-id-ariacovers ids something actually references, is still live WCAG 4.1.2, and needs no enabling — but axe marks itreviewOnFail, so its findings arrive asincompleterather thanviolations. That left the worst case invisible: two<input id="q1">under one<label for="q1">returned zero violations even with both obsolete rules on. A duplicate id needs no human judgement to confirm, so this rule's incomplete results are promoted to violations — only this rule, since the rest ofincompletegenuinely cannot be decided without rendering.
This widens what the gate reports, which is the point but has a cost worth knowing: a document that used to pass now spends review iterations on duplicate ids, and can reach
max_review_iterationswith them still listed inunresolved.md. Assembly namespaces the cross-page duplicates itself, so what reaches the review loop is the ids duplicated within a single page — which the assembler cannot fix, because there is no second page to attribute the copy to — plus the collisions on any page the reserialization guard left as written. -
Colliding ids are namespaced during assembly (§7.7 v1.2). A page is extracted alone and concurrently, so it cannot know that another page also numbered its first footnote 1 — and the page prompt asks it to preserve the source numbering.
assembleBodyprefixes the ids that more than one page claimed with their page number (fn-1→p3-fn-1) and rewrites everything that points at them in the same pass:href="#…", plusfor,headers,list,formand thearia-*references, since unique ids with dangling references would be a worse defect than the collision.The scope is deliberately one id at a time, not one page at a time. Prefixing every id on a page also breaks the references that legitimately span a page break — a
<label for>whose input is on the next page, or endnotes with continuous numbering — which resolved correctly before assembly touched them, so that trade is a no-target reference in place of a wrong-target one.The prefix is reserved against every id the document already claims, growing its separator (
p1-→p1--→ …) until nothing collides with it, becausep1-totalandp2-nameare what a paginated form emits and a blind prefix would manufacture the duplicate it exists to remove. An ordinary document keeps the short form.The prefix is labelled with the page number, but it does not depend on that number being unique: two fragments sharing an
orderwould otherwise take the same prefix and stay collided, with the log reporting the id as namespaced. Ownership is tracked per fragment position and a repeated label becomesp1_2-.Every reference to a colliding id is repointed rather than abandoned. If the page owns the id it goes to the page's own copy (reference and target were written together by one agent looking at one image). If it does not, the reference is ambiguous and goes to the first page in document order that claims the id — where a browser sent the bare reference before any of this ran. Leaving it dangling instead was the same defect in a new place: with a
<label for="q1">on page 1 and an<input id="q1">on pages 2 and 3, every owner is renamed and the label points at nothing, so the field loses its accessible name and axe reportslabelon a document a plain concatenation passed. Ambiguous references are named in the run log asassembly_anchors. A page whose markup would not survive a reserialization is left exactly as written, keeping its collision for lint to report and its bare ids for anything resolved to it. If such a page holds a reference instead, the referenced id's first owner keeps its bare form so that reference still resolves — only the first owner, so every other copy is still renamed, and only when none of that id's owners was skipped, since a skipped owner is already keeping the bare id and pinning a second copy would ship a duplicate. Any id pinned this way is listed in the same log line aspinned_ids: it is a colliding id that deliberately was not renamed, so without it a bare colliding id in the delivered document would be indistinguishable from namespacing that silently failed. A page too deeply nested to rewrite — rewriting recurses per level in three places, so past 500 levels, measured on the parsed tree, the page is refused rather than allowed to overflow one of them — is delivered as written for the same reason and takes the same treatment: it counts as an owner (or the collision would go undetected for its copy, and the pin would fire on top of the bare id it is already keeping) and its frozen references pin their first owner. Its ids and references are read from its DOM, which such a page keeps:querySelectorAlldoes not recurse, so it works at any depth the parse survived, and the reading is exact. Only a page whose parse threw falls back to scanning the source, and that scan follows the parser's own rules — attributes only from real tag positions, elements whose content is not markup (<textarea>,<script>,<template>and the rest) skipped, character references decoded, first of a repeated attribute — because a phantom id read out of non-markup text is worse than a missed one: it suppresses the pin, the real owner is renamed, and a<label for>elsewhere is left naming nothing. Reading the tree is what closed that class rather than modelling more of the parser: the scan cannot see tree construction, so it invented owners for markup the parser drops outright (an orphan<tr>/<td>, a stray<caption>/<col>/<thead>, anything after<plaintext>) and missed real references inside a<select>, whose<option>children survive parsing even though most tags in there do not. That covers foster parenting in both directions: a<tr>outside a<table>is dropped to bare text, and content inside one is hoisted out past the table — a reading-order change, worse than the duplicate id it would be fixing. The guard compares the source's sequence of tags and text against the parsed document as a subsequence, since counts cannot see a move, equality would refuse every page where the parser legitimately adds a tag, and a tag-only sequence misses bare prose being hoisted out of a table with every tag left in place. -
Copy Editor image payload (§7.9). When every issue in a round is attributed to a page, the editor gets only those pages' images (logged per round as
editor_images). Attaching every page's image on every round is the dominant per-round cost of the review loop — on a 25-page document that is 25 base64 PNGs × up tomax_review_iterations. Narrowing requires full attribution: one unattributed issue re-broadens the round to every image. An unattributed issue is usually structural and fixable from the HTML alone, but it is also what a heavily editor-rewritten body looks like once it no longer matches the source excerpts — so narrowing wrongly can leave a real issue unfixed at the iteration cap, while broadening wrongly costs no more than the behavior this optimization replaced. -
The flattened screen-reader view must never lose text (§7.8).
flatten.tshas two consumers, and both fail silently when text goes missing: the Reader reviews this view instead of the source images, so anything absent from it cannot be reported as an issue; andcontentCoveragemeasures a candidate agent against an accepted fixture using these words, so text the view can't see is absent from both sides of the comparison. The second is the sharp edge — the regression gate exists to stop an agent update from dropping content, and it scored a table whose every row had been deleted as perfect, because the old implementation emitted a table's<caption>and returned. Inline elements (a,img,em, …) are now announced within the surrounding phrase and block elements are separate stops, with tables expanded row by row;test/flatten.test.tsasserts the invariant mechanically by deriving the expected word set from the DOM independently offlatten. Both halves of that inline/block split recurse, so the same pathological nesting the assembler delivers rather than drops would overflow the stack here and throw — losing all the text, the worst form of the failure. The walk therefore falls back to an iterative pass that keeps words and reading order and gives up structure, which is the trade the view already makes for a block inside a table cell. Role markers are stripped before the coverage comparison anyway, so a marker-free view scores identically while a dropped word still registers.Two rules follow from
contentCoveragestripping[...]before it compares words, and both are easy to break by accident. Everythingflattenadds itself must be inside brackets — including annotations that read like prose ([3 rows, 2 columns],[empty],[spans 3 columns],[alt missing]) and a control'stype, which a screen reader announces as its role. An unbracketed annotation is counted as a word the agent produced and is reproduced free by any candidate emitting a similar structure, which pads the ratio:(2 rows, 3 columns)alone moved a fixture that had dropped a table row from a true 0.833 to a reported 0.875, across the 0.85 gate. And a field's text lives in its attributes, not its child nodes — so every code path must announce fields through the one shared helper. When only the block path did, a field inside a table cell or an inline wrapper contributed nothing and a form-as-table with every value emptied scored 1.0.test/flatten.test.tsenforces the first rule generically (nothing outside brackets may be a word the source document doesn't contain) rather than by listing known markers, which is what let the parenthesised ones slip through initially.A third rule, learned the same way: an accessible name can live in an attribute (
aria-label,title), so those count as announced content — an agent update that dropped everyaria-labelscored 1.0 before and 0.3 after. The test baseline deliberately collects a wider attribute set thanflattenreads, because when the two lists matched the baseline shared the code's blind spot and no attribute loss could fail a test. A baseline derived from what the code looks at is not independent of the code.The prompt and the markers are one contract in the other direction too:
test/flatten.test.tsassertsREADER_SYSTEMadvertises no markerflattennever emits ([Option]was documented and unreachable), and every annotation that explains correct markup —[spans N columns],[spans N rows],[decorative, alt empty]— exists because the prompt tells the Reader that an unexplained mismatch is a defect, and the Copy Editor is licensed to restructure tables. Adding a check to that prompt without the annotation that reconciles it turns the review loop into a false-positive generator aimed at accessible output. -
Both sides of the eval gate must score fixtures by the same rule (§7.12). Before proposing an agent update, Iris compares the candidate prompt's mean fixture coverage (from
regressionGate) against the current prompt's (fromevalAgent) and blocks a drop of more thanEVAL_REGRESSION_EPS(0.02). That comparison is a subtraction between two means, so it is only valid if both are computed identically — and they were not.contentCoveragereturnsnullfor a fixture whose accepted text is underMIN_COVERAGE_WORDS(8) because one dropped word would swing the ratio;regressionGateexcluded those from its mean, whileevalAgentscored them a perfect 1. Since abstention depends only onaccepted_html, the same fixture abstained on both sides, so the 1 landed on the current-prompt side alone and inflated it. WithMAX_GATE_FIXTURES= 3 that is large: two judgeable fixtures at 0.90 plus one unjudgeable gave current 0.933 vs candidate 0.900 — a 0.033 gap from padding alone, past the 0.02 threshold. The gate discarded updates whose measurable coverage was identical, logged aseval_regression: a reason naming a regression that had not happened. A singlefixtureScorehelper now defines the rule for both, and an abstaining fixture is absent from both sides rather than scored. Note the direction — the failure mode here is a false block, not a wave-through, which is why it was invisible: a learning loop that silently declines to learn looks like a loop with nothing to learn. A mean over zero measurements isnull, not 0 — the caller treats that as "nothing to compare" and defers to the regression gate, since 0 would block every update and 1 would assert a score no fixture demonstrated.No output at all is scored 0 rather than abstaining, because producing nothing is a failure on the fixture, not an absence of evidence — abstaining would let a prompt that returns nothing score as well as one that handles it. That is also the one input where abstention is not purely a property of the fixture: whether a prompt produced output is a property of that prompt, so one fixture can be scored 0 for one side and excluded from the other.
-
The eval gate is a paired comparison, per fixture (§7.12). The rule above is right about what a score means, but averaging each side over whatever it happened to measure compared two different fixture sets — and in one direction that waved a real regression through. If the current prompt flaked to no output on a fixture the candidate abstained on, the current mean was deflated and the bar dropped: one such fixture plus one judgeable at 0.98 gave current
(0 + 0.98)/2 = 0.49against a candidate at 0.88, so0.88 < 0.49 - 0.02was false, 0.88 cleared the 0.85 floor, and a real 0.10 coverage regression passed both gates. Note this is the opposite direction from the false block above — the same asymmetry, read from the other side.Both scorers now return per-fixture scores and
pairedMeansaverages only the fixtures both prompts could be scored on, so a per-prompt exclusion drops the fixture from both means instead of moving the threshold. Deliberately, a current-prompt flake is treated as evidence for neither side: it is a problem with the current library agent, and lowering the bar is the one response that hides both it and any regression behind it. It stays visible in theeval_gatelog line'sunpairedlist. If no fixture is measurable on both sides, both means arenull— "nothing to compare", deferring to the regression gate, rather than a pass. -
GET /v1/sessionspages on a compound cursor (§9.2 v1.1). The PRD names acursorparameter without saying what is in it, and the obvious reading — the last row'screated_at— is unsound:created_atis a millisecond timestamp assigned by a request handler, so a burst of uploads ties on it, and paging on a non-unique key skips rows (created_at < ?drops the rest of a tied group) and can repeat them (nothing pins the order among ties).next_cursoris therefore"<created_at>|<session_id>", the full sort key; clients pass it back verbatim. A cursor that doesn't parse is a400, not a silent restart at page one, andnext_cursorisnullon a full final page — so clients stop on a null cursor rather than on a short page. -
Runs are queued, and the queue is in-process (§9.4). A bounded FIFO queue (
src/util/queue.ts) caps concurrent pipelines atdefaults.max_concurrent_runs; sessions over the cap wait inqueued. Two things this deliberately does not do. It does not persist: the queue lives in the process, so a restart loses waiting runs — they are markedfailed("interrupted (server restarted)") by the samefailStaleSessions()sweep that already handled interruptedrunningsessions, which is why that sweep coversqueuedtoo. And it does not bound upload memory: multer parses the whole body before any handler runs, so by the time the queue sees a session its images are already buffered in RAM (ceiling: multer's ownlimits.fileSize× part count) and any PDF is already rasterized to full-page 150-DPI PNGs. Both are consequences of the single-instance, single-process design the store declares. -
Starting work on a session is a claim, not a check (
store.claimSession). The two endpoints that begin non-idempotent work —POST /:id/feedback(enqueues a pipeline) andPOST /:id/close(files regression fixtures into the shared agent library, deletes the tmp tree) — used to read the status, compare it, then write.claimSessionfolds the comparison into the write (UPDATE … WHERE session_id = ? AND status = ?) and reports whether this caller is the one that changed the row, so of two concurrent callers exactly one is told it won.What this is and is not: both handlers are fully synchronous, so today nothing can interleave between the check and the write and the plain pattern was already correct. Racing two processes against a shared WAL database, both callers won — but a second instance is not the supported topology (see the in-process queue above). So this is defense in depth. It earns its place by being the cheaper invariant to hold: correctness stops depending on every future handler staying synchronous. Adding one
awaitbetween the guard and the write — the ordinary thing to do when a check needs I/O — would silently reintroduce the race in-process, and a duplicated feedback run is invisible in the response (both callers get a202) while two pipelines write the sameoutput.htmlandfragments/final.json.The claim sits last in the feedback handler (after request validation, so a malformed body still gets its
400without disturbing the session) and first in close (before fixture capture and thermSync, because a loser that discovers it lost afterwards has already filed the fixtures twice). -
Provider retries are not symmetric in code, but are in behavior. OpenRouter retries by hand (3 attempts, exponential backoff) because
fetch()has no retry strategy. Bedrock has no retry loop on purpose: the AWS SDK already applies itsstandardstrategy — also 3 attempts with exponential backoff — to throttling, 5xx, and node network errors, while failing fast on 4xx. Verified empirically against a stubbed request handler (3 wire attempts for 503/429/ECONNRESET, 1 for a 400). Adding a loop around it would give Bedrock 9 attempts to OpenRouter's 3. -
Feedback re-runs (§7.12). Re-runs are logged separately (a
feedback_rerunevent) and the prioroutput.htmlis snapshotted tosessions/<id>/history/so it can be reverted to. A revert endpoint is out of v1 API scope (not in §9); the data is preserved to enable it.A re-run is routed first (
feedback_scopedevent). The Reader only ever sees the assembled HTML (by design, §7.8), so feedback about what was read off a page ("the revenue figure on page 2 is wrong") raises no issue for the loop to act on and cannot be fixed there. The Feedback Agent's SCOPE task decides which case applies:document— tone, wording, ordering, or an accessibility rule: re-lint the saved body and run the feedback-aware review loop on it. No source images, no re-extraction.extraction— source-fidelity: the named pages go back to the page agent with their source image and their previous output attached, then the document is reassembled and reviewed. Untargeted pages keep their prior fragments byte-for-byte.
Routing is deliberately biased toward the cheap path: an unavailable agent, an unparseable answer, pages it cannot localize, or a claim spanning more than half the document all fall back to
document. A wrongdocumentanswer costs one review round; a wrongextractionanswer costs a vision call per page. -
One instance per
data_dir— this is a hard constraint, not a preference. Running two processes against the samestorage.data_dircorrupts sessions, and it fails loudly in the wrong direction: on boot each instance runsfailStaleSessions(), which marks everyrunningandqueuedrowfailedwithinterrupted (server restarted). Those rows include the other instance's live runs. A second instance starting therefore kills the first one's in-flight conversions from the client's point of view — the pipeline keeps going and still writesoutput.html, but the session readsfailed, so the user is told their document failed while work continues on it. The sweep cannot tell "this row is orphaned" from "this row belongs to a peer" because nothing records which process owns a run.Two other single-process assumptions ride along: the run queue that enforces
max_concurrent_runsis in-memory, so N instances allow N × the cap, and fixture and agent-memory writes underdata_dirare unsynchronized between processes.To scale beyond one box, put a second
data_dirbehind it (independent instances, sessions not shared) rather than pointing two at one directory. Gating the sweep on an instance id, and moving the queue and locks out of process, is what a genuinely multi-instance version needs. -
phasereports only phases that exist.extraction,assembly,review,done. The PRD'striage(§7.2) andreconciliation(§7.6) are not implemented — reconciliation is unreachable while extraction hardcodesedges: []— so they are not in the enum and not emitted (§9.2 v1.1). New sessions start atextraction; they used to be created attriageand overwritten before a client could observe it.
Intentionally not built in v1 (the PRD frames each as optional / alternative / out of scope):
PostgreSQL and S3 backends (§10.2 — "supported alternative," SQLite + local FS is the v1
reference), the per-user config endpoint (§9.1 — "not specified in v1"), and webhooks (§9.4 —
out of scope). The only endpoint beyond the PRD is GET /v1/health, a standard liveness probe.
Every PR is reviewed by Claude in CI before a human reads it
(.github/workflows/code-review.yml, PRD §7.14). This is
not convenience tooling. Iris's agent library only improves through upstream merge (§7.13), so
review capacity is the bottleneck on the whole contribution model — and a three-institution
maintainership with no full-time reviewer cannot be the only thing between a contributed prompt
and every future session.
What it does, in order:
- Runs
npm ci,tsc --noEmit, the unit suite,./test/e2e.sh, andactionlint, and hands the model their actual output. The reviewer is told not to re-run them, so a claim that a check failed is quoted rather than predicted. - Builds a context file: the diff, plus full source for files that are new or substantially rewritten, plus up to the 3 most recent prior reviews on earlier commits of the same PR — so a re-review knows what it already said instead of repeating it.
- Reviews against a ranked list: accessibility of the output, upstream side effects and filing identity, auth/tokens/secrets, provider routing and cost, correctness, failing checks, missing tests, and the PR template's own contract.
- Posts exactly one review ending with a one-line
Accessibility impact:.
Blocking is decided by reachability, not by category. A finding blocks only if a real user, a
real request, or CI reaches it on input the code accepts today, and each blocking finding has to
name that input. A defect that's real but unreachable is a note on an approval, with what
would have to change to reach it. This was tuned in response to a measured problem: the findings
were reproduced and specific, but everything arrived as blocking — 34 CHANGES_REQUESTED to 17
APPROVED across the repo's history, individual PRs at 12-to-1, including reviews that called
their own finding latent and requested changes anyway. main has no branch protection, so the
cost was never blocked merges; it was author attention, and a reviewer that always blocks trains
you to skim the one time it matters. Three things stay blocking even when unreachable, because
their value is holding when something else breaks: auth/token/secret handling, publishing under
the wrong identity, and path handling that could escape the data dir.
Depth was not what got trimmed. The model gets ~19 minutes and is told to dig exactly as hard
as before; the bar governs the verdict, not the investigation. It appends findings as it confirms
them, so if it's cut off, a fallback step posts the partial findings plus the check summary as a
--request-changes — an incomplete review must not read as a pass. A final step fails the job if
no review was posted at all, since the action can exit 0 without posting one.
Two gaps worth knowing:
- A PR that modifies
code-review.ymlgets no automated review.claude-code-actionrefuses to run when its own workflow file differs from the copy onmain. Everything else on such a PR still goes green, so the workflow says so loudly — a step-summary block and an Actions warning naming what to check by hand (does it let PR-authored code run with secrets, widenpermissions:, or interpolategithub.event.*into arun:block). A second workflow used to cover this onpull_request_target; it never produced a review in six runs and was the repo's only PR-triggered job holdingid-token: write, so it was deleted and the gap accepted. - Fork PRs are skipped.
pull_requestfrom a fork gets no secrets, so the OIDC role assumption would fail confusingly. Review one withgh workflow run code-review.yml -f pr_number=<n>— which runs the fork's code in a job holding the Bedrock role, so read the diff first.
The verdict is advisory: main is unprotected and a human still merges. What changes is what
that human is reading, not whether they read it.
See CONTRIBUTING.md and our Code of Conduct. Found an accessibility barrier — in the app or in the HTML it produces? Please open an Accessibility issue; those are our top priority.
PRs get an automated review before a human reads them — see Automated code review above for what it looks at and, more usefully, what it deliberately does not flag (style, formatting, naming, "you could also do X", pre-existing issues your PR doesn't touch).
GNU AGPL-3.0-or-later. Iris is copyleft: if you modify it and run it as a network service, you must make your modified source available to its users (AGPL §13).
Iris is maintained by Equalify Inc., the University of Illinois Chicago, and California State University.
Commercial hosting and support are offered by Equalify Inc. The hosted and self-hosted versions are functionally identical — what you are paying for is operational (managed deployment, monitoring, accessibility consulting), not features withheld from this repo. Please consider hiring them to host or support your instance.