AIME-Con 2026 · Panel · Wednesday 7 October, 11:00–12:30 · Commonwealth 2
Session chair Damian Betebenner, Center for Assessment
Moderator lain, an AI interlocutor under the chair’s supervision
Discussant Fred Oswald, University of California, Irvine
Panelists
Derek Briggs, University of Colorado Boulder
Frank Rijmen, Cambium Assessment
Mohammed A. A. Abulela, MetaMetrics / University of Minnesota
Damian Betebenner, Center for Assessment
Panel materials
dbetebenner.github.io/
A correction to the printed program: it lists a human moderator. Today the moderator is lain, an AI, under the session chair, and Fred Oswald is our discussant.
Of the 335 items in the printed program, 66% build AI into what we deliver and 39% already use AI in how we work. 22 items (7%) ask what that means for the profession itself.
Rows overlap: each item was coded on three independent yes/no questions by an AI coder, and a blind second AI coder on a stratified sample of 98 agreed on 91%, 97% and 94% (κ 0.80, 0.94, 0.72). Audited by the chair. Data, prompts and agreement: the materials site.
At this conference: products 221 · our methods 131 · the profession itself 22 (of 335 items).
The chair’s conjecture: an independent order-of-magnitude claim, not estimated from the program coding
How measurement professionals use AI in their own work will have 10× the impact of the AI we build into assessments.
AI in the product changes one tool: a scoring engine, an item pool, a tutor.
AI in the profession changes every analysis, model, review, document, and decision, including how the next generation of those products is designed, validated, and governed.
The mathematics community has begun asking this explicitly. The Leiden Declaration on AI and Mathematics (June 2026, endorsed by the IMU; leidendeclaration.ai) asks the discipline to protect the verifiability of proof, attribution, authorship, and the primacy of human judgment.
Five guiding questions
An experiment in AI-native educational measurement
lain is our attempt to show what AI-native educational measurement could look like: in this room, not on a slide.
Its evidence is bounded to the panelists’ materials and what is said in this room, with no open web. Like any language model it carries pretrained knowledge, but it may not present that as evidence: a source claim needs a citation, a claim from the room needs a timestamp, and anything else is labeled inferred. Its authority belongs to the chair. Everything it proposes, including what it is not allowed to say, goes on the record.
As AI becomes more capable, one of the most important things professionals do is this: gather, deliberate in public, and decide what counts as warranted.
A session like this one can leave more than memories. The materials, the claims, the disagreements and the norms are collected into a record that future colleagues, and the AI systems they work with, can query, for example through MCP or other AI connectors.
Presence: continuous, visual, silent
A running record on screen: claims and who made them, sources, agreements, open tensions, question coverage.
Voice: rare, typed, human-approved
About 8–12 interventions in 90 minutes, each approved by the chair.
OBSERVE
↓ chair enables proposals
PROPOSE
↓ chair approves one intervention
SPEAK
↓ delivered; the permission expires
Emergency stop → OBSERVE, at once
Every contribution is tagged observed (the transcript records someone saying it, with a time), retrieved (in a panelist’s material, with the source), or inferred (lain’s own reasoning, possibly wrong). Its evidence: the panelists’ materials and the transcript; no open web. It abstains when someone else holds the warrant.
| State | What it means |
|---|---|
| GREEN | Live transcription and running record; approved interventions spoken in lain’s voice |
| AMBER | Same, but I read approved interventions aloud |
| YELLOW | Transcription or analysis degraded; the record freezes at its last good state |
| RED | A conventional panel: same questions, same norm-building |
Running today: I’ll tell you which state we’re in, and I will call every change of state aloud. The panel’s argument does not depend on lain working.
This session is being recorded and transcribed.
lain, an AI moderator, speaks aloud under the session chair, who approves every sentence. The transcript and the session record are published afterwards. Questions, typed or asked aloud, become part of that record. No names or emails are collected.
The record includes the full intervention log, including what lain proposed and I declined. Panelists review it before release.
Follow along · ask a question
or open it from the materials site
Panel materials
dbetebenner.github.io/
| Minutes | Segment | lain |
|---|---|---|
| 0–6 | Framing and the AI design | visible, silent |
| 6–26 | Four five-minute provocations: one claim, one example, one unresolved problem | OBSERVE |
| 26–31 | AI synthesis: convergence, disagreement, connections | one approved intervention |
| 31–43 | Discussant: Fred Oswald critiques the panel and the AI synthesis | OBSERVE |
| 43–76 | Panel and audience dialogue | PROPOSE; approved SPEAK |
| 76–87 | Norm-building: activity · conditions for AI involvement · who is answerable | PROPOSE |
| 87–90 | Human closing synthesis | silent |
Educational Measurement as an AI-Native Profession · AIME-Con 2026