# Audit sheet: two-dimension recode (6 Oct 2026)

## 1. The 11 items where the two coders disagree (n = 98 sample)

| id | title | coder 1 | coder 2 | coder 1 rationale | coder 2 rationale |
|---|---|---|---|---|---|
| P008 | What Are We Assessing in Human-AI Interaction? (Keynote: Dragan Gaševi | product only | neither | Assessing learning when performance is produced through human-AI interaction in tasks. | Assessing human learning amid AI use; AI not product nor professional tool |
| P036 | Rethinking Pilot Data: Evaluating LLM Synthetic Data for Scale Develop | both | profession only | LLM synthetic respondents for early psychometric analysis in scale development. | LLM synthetic data for psychometric scale development analysis |
| P073 | Using Generative AI to Simulate Item Responses by Skills Insight Abili | both | profession only | GPT simulating item responses by ability band for early item evaluation. | GenAI simulated responses for early item evaluation before field testing |
| P138 | Introduction to AI for Measurement | both | profession only | Title only: introduction to AI for measurement. | Title only: introduction to AI for measurement |
| P143 | Building Responsible AI Practice in Language Assessment: Anticipating  | both | profession only | Title only: responsible AI practice in language assessment. | Title only: responsible AI practice in language assessment |
| P155 | Human–AI Collaboration in Educational Measurement: Transforming Assess | both | profession only | Title only: human-AI collaboration in measurement turning assessment data into action. | Title only: human-AI collaboration in measurement work |
| P233 | Measuring Collaborative Reasoning with LLMs | product only | profession only | LLMs classifying collaborative reasoning in student discussions. | LLMs labeling student discussion reasoning; coding of discourse data |
| P257 | S2A3: Thompson Sampling and Stochastic Exposure Control for High-Stake | product only | neither | Bayesian Thompson sampling adaptive testing with continuous calibration. | Bayesian adaptive testing and exposure control; no clear AI |
| P282 | The Analysis of Artificial Intelligence’s Decade-Long Impact on Educat | both | profession only | Content analysis of AI's decade-long evolution in educational measurement. | Synthesis of AI's impact on the measurement field |
| P283 | Transformer-Aided Detection of Gaming in Constructed-Response English  | product only | both | Transformer detection of gaming in automatically scored responses. | Transformer detection of gaming in constructed responses; scoring and QC |
| P304 | SarphieSense: A Pipeline-Based AI Chatbot for Navigating Living System | neither | profession only | Title only: AI chatbot for systematic review evidence; not assessment or measurement work. | Title only: chatbot navigating systematic review evidence for researchers |

## 2. Old 'product' (A) items now coded profession only (coder 1)

| id | title | rationale |
|---|---|---|
| P010 | AI-Enabled Quality Assurance for Multiple-Choice Assessment Items | AI-enabled quality assurance of multiple-choice items: AI as item reviewer. |
| P011 | Predicting Item-to-RPLD Matches with Structured AI Reasoning | LLMs classifying items to performance level descriptors to support assessment design. |
| P012 | Can AI Detect Item-Metadata (Mis)alignment? An Empirical Study Using State Asses | LLMs classifying items to standards for assessment quality assurance. |
| P013 | Achievement-Level Descriptor Conditioned Option-Selection Probability Modeling w | LLM-supported item evaluation predicting option selection and item difficulty. |
| P038 | Prior-Informed 3PL Calibration: Reducing Sample Size via Predictive Modeling | Predictive item-parameter priors to reduce calibration samples; AI implied by session. |
| P051 | Putting it Together: Using Embeddings Clusters to Identify Inconsistent Raters | Title only: embedding clusters to flag inconsistent human raters; operational quality control. |
| P056 | Exploring AI-driven Methods for Pre-calibrating Difficulty of Mathematics Items  | AI pre-calibration of item difficulty replacing field-testing labor. |
| P058 | What LLMs Capture and Miss About Cognitive Sources of L2 Item Difficulty | LLM zero-shot item difficulty estimates versus cognitive attributes. |
| P059 | Math Item Difficulty Prediction with Multimodal Input | Multimodal LLMs predicting math item difficulty. |
| P060 | Beyond Item Text: Modeling Latent Cognitive Demand for Math Item Parameter Model | Modeling cognitive demand to improve item parameter prediction. |
| P087 | LLM-Based Pairwise Judgment for Math Item Parameter Modeling using Workflows and | LLM pairwise judgments to recover item parameter scales. |
| P088 | LLM as Investigator and Judge in Pairwise Comparisons for Item Parameter Modelin | LLM investigator and judge for item parameter modeling. |
| P089 | Validating LLM Rating of Item Features for Explanatory Item Response Model . | LLM rating of item features for explanatory IRT difficulty prediction. |
| P107 | Partial Identification with Multiple Nonlinear Measurements of a Latent Regresso | Estimator combining disagreeing AI scores for downstream research analyses. |
| P108 | Demonstrating FairGAIte: An Agentic LLM Tool for Detecting Construct-Irrelevant  | Agentic LLM tool for fairness review of pilot items. |
| P110 | Using LLM Judges’ Paired Comparisons to Estimate Mathematics Item Difficulty | LLM paired comparisons to estimate item difficulty without field testing. |
| P111 | Evaluating Multiple Models for Predicting Item Difficulty in a Principled Assess | ML models predicting item difficulty in principled assessment design. |
| P112 | Consensus without Accuracy: Investigating LLMs Recovery of Item Difficulty Using | LLM pairwise comparisons recovering item difficulty for AI-assisted calibration. |
| P113 | When One Benchmark Hides Many Truths: Mixture IRT and LLM Difficulty Prediction | Validating LLM difficulty prediction against mixture IRT classes. |
| P117 | A Multimodal AI Analysis of Teacher Praise in Online Tutoring | Multimodal AI coding teacher praise quality in tutoring research data. |
| P118 | Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions | AI transcription of tutoring sessions to build research logs. |
| P125 | AI-Assisted Rater Training Modules: Effects on Qualification and Scoring Quality | Title only: AI-assisted training modules for human raters. |
| P128 | Using Natural Language Processing to Explore Alignment Between Skill Taxonomies | NLP analysis of alignment across skill taxonomies. |
| P129 | Pre-Data Embedding Diagnostics for Construct Validity in Multi-Domain Instrument | Embedding diagnostics for construct boundaries in instrument development. |
| P130 | Auxiliary Information for Semantic Clustering of Assessment Items | Transformer clustering of operational items for blueprint and SME review. |
| P131 | Supporting Distractor Quality Review Through Interpretable Semantic and Lexical  | NLP diagnostic tool supporting distractor quality review. |
| P135 | Automatic Pre-Testing of Mathematics Items: predicting difficulty parameter with | ML automated pretesting to predict IRT difficulty for calibration. |
| P137 | Using XGBoost to Construct a Vertically Skill Difficulty Scale | XGBoost model building a vertical skill difficulty scale. |
| P145 | Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assess | Factor analysis comparing human and LLM respondents interpreted by experts. |
| P149 | Predicting IRT Parameters for Passage-Based Reading Items Using NLP | NLP features predicting IRT parameters of reading items. |
| P150 | Predicting Item Statistics Using LLM-Derived Difficulty Ratings and Response-Opt | LLM-derived features predicting item statistics for certification items. |
| P151 | Anchored Bradley-Terry Calibration Using LLM Comparative Judgments | LLM comparative judgments for early item difficulty calibration. |
| P156 | Understanding Spatial Cognitive Process Using Time-Embedded N-Grams Model with M | Title only: machine learning analysis of spatial cognitive process data. |
| P165 | Hierarchical Attention Network Architecture for Automated Response Time Predicti | Deep learning predicting item response times from metadata. |
| P168 | Fine-tuning DeBERTa-v3 to Automate Spatial Language Classification | Fine-tuned DeBERTa replacing human coders for spatial language classification. |
| P187 | A Specialized Multimodal Transformer for Classroom Discourse Classification | Multimodal transformer classifying classroom discourse data. |
| P192 | Measurement-Informed Difficulty Priors for Competition Math Items | LLM features for item difficulty priors. |
| P197 | Lost Without Translation? Multilingual Sentence Embeddings for Linguistic-integr | Multilingual embeddings replacing translation in scoring reliability auditing. |
| P227 | AI-Assisted Evidence-Centered Design to Build a Blueprint for a Semiconductor Cr | Title only: AI-assisted ECD to build a credential blueprint. |
| P234 | An exploration of using an AI-based approach to shorten an assessment | AI shortening an assessment compared with expert-shortened form. |
| P244 | Does including images improve multimodal language models’ accuracy on item discr | Multimodal models predicting item discrimination. |
| P245 | Image-Based Representation of Item Response Patterns for Test integrity | Deep learning anomaly detection for test integrity. |
| P246 | Listening for Meaning: Evaluating an ASR-LLM Pipeline for Thematic Audio Analysi | ASR-LLM pipeline coding student audio themes to inform assessment development. |
| P265 | Beyond Latent Factors: Machine Learning Approaches to Response Clustering | ML clustering of survey responses as alternative to factor analysis. |
| P277 | Why AI Demands More Principled Assessment Design, Not Less | Title only: principled design practice for AI-assisted assessment design. |
| P279 | AI-Generated Distractor Plausibility Ratings as Predictors of Mathematics Item D | Title only: AI distractor plausibility ratings predicting item difficulty. |
| P280 | Predicting Item Difficulty: How PAD Features Shape Model Performance | Title only: PAD features in item difficulty prediction models. |
| P295 | Design Choices in Encoder-Based IRT-3PL Parameter Prediction for Passage-Based M | Encoder-based IRT parameter prediction for item calibration workflows. |
| P300 | Automated Item Evaluation: Predicting Item Acceptance and Rejection using LLM-Ge | Transformer models predicting operational item acceptance from critiques. |
| P303 | EquiFrame: An AI-based Program that Assesses Disability Language in Educational  | Title only: AI tool reviewing disability language in measurement instruments. |
| P309 | From Validation to Resilience: AI-assisted Enemy Item Identification in Operatio | Title only: AI-assisted enemy item identification in operations. |
| P311 | Predicting Race and Ethnicity DIF Using Generative AI in Large-Scale Assessment | Generative AI DIF screening supporting item review in test development. |
| P315 | Fairness Sentinel: An AI Agent for Early DIF Risk Screening | AI agent screening draft items for DIF risk before field testing. |
| P324 | Offloading the White Elephant: Assembling with AI-estimated IRT Parameters Inste | AI-estimated IRT parameters replacing pretesting in form assembly. |

## 3. The 6 norms disagreements (n = 98 sample)

| id | title | coder 1 | coder 2 | coder 1 rationale | coder 2 rationale |
|---|---|---|---|---|---|
| P154 | To Err is not only Human: Human/LLM Collaboration in Assessment | 1 | 0 | Panel on human/LLM collaboration and error in assessment work; implies division of verification responsibilities. | Human/LLM collaboration panel; no abstract |
| P173 | AI-Augmented Form Assembly from a Secure Item Bank: A Faculty Pilot | 1 | 0 | Faculty test builders using AI; examines their judgment boundaries in AI-assisted form assembly. | Faculty AI-assisted form assembly pilot; judgment boundaries tangential |
| P177 | Building Teacher Capacity to Generate and Evaluate Math Items with Sma | 1 | 0 | Professional development model building teacher capacity to generate and evaluate items with LLMs. | Teacher PD for AI item generation; teachers not measurement professionals |
| P277 | Why AI Demands More Principled Assessment Design, Not Less | 0 | 1 | Argues AI requires more principled assessment design; design methodology rather than professional norms. | Argues AI demands more principled assessment design practice |
| P282 | The Analysis of Artificial Intelligence’s Decade-Long Impact on Educat | 0 | 1 | Bibliometric synthesis of AI research trends in measurement; field content, not norms or roles. | Charts AI's decade-long impact on the measurement field; responsible adoption |
| P335 | PsyMAS: A Human-in-the-Loop Multi-Agent System for Auditable Test Secu | 0 | 1 | Human-in-the-loop test security system keeping misconduct decisions human; system design foremost. | Analyst adjudication, auditability; AI must not make misconduct decisions |

## 4. The 22 items coded norms = 1 (coder 1)

| id | title | rationale |
|---|---|---|
| P004 | Integrating Generative AI into R Workflows: From APIs to Shiny Apps | Frames a training gap for measurement professionals expected to adopt AI while upholding standards. |
| P007 | Writing an AI-Native Dissertation | Reimagines doctoral training: building an AI-native dissertation, a change in how researchers are formed. |
| P033 | Industry Leader Perspectives: How AI is Changing K12 Measurement Roles | Panel explicitly about how AI is changing K12 measurement roles. |
| P046 | Responsible use of generative AI when creating reading comprehension questions:  | Argues responsible-use practice: AI item evaluations should document prompts and generation steps. |
| P085 | Our Future with AI: Graduate Student Perspectives on a Changing Profession | Panel on how AI changes the profession, graduate training and career preparation. |
| P086 | A Jurisdictional Scan of Automated Scoring Practices in K-12 Summative Assessmen | Scan of automated-scoring practices across state programs; describes operational practice landscape of the field. |
| P109 | Human-Centered AI Applications: Bridging Learning Science, Assessment, and Ethic | Panel on governance choices sharing responsibility and authority between humans and AI in assessment. |
| P125 | AI-Assisted Rater Training Modules: Effects on Qualification and Scoring Quality | AI-assisted rater training and qualification; subject is training of the scoring workforce. |
| P139 | Human agency before and after Generative Artificial Intelligence in education, a | Human agency in measurement before and after GenAI; plausibly addresses professionals' authority and role. |
| P143 | Building Responsible AI Practice in Language Assessment: Anticipating What’s Nex | Panel on building responsible AI practice in language assessment; professional norms. |
| P154 | To Err is not only Human: Human/LLM Collaboration in Assessment | Panel on human/LLM collaboration and error in assessment work; implies division of verification responsibilities. |
| P173 | AI-Augmented Form Assembly from a Secure Item Bank: A Faculty Pilot | Faculty test builders using AI; examines their judgment boundaries in AI-assisted form assembly. |
| P177 | Building Teacher Capacity to Generate and Evaluate Math Items with Small LLMs | Professional development model building teacher capacity to generate and evaluate items with LLMs. |
| P189 | Interweaving Assessment: A Field-Generated R&D Agenda for AI and Measurement | Field-generated R&D agenda for AI and measurement; the field setting its own direction. |
| P221 | Implementing Generative AI Assessment Guidelines in a Mexican University | University policy process for GenAI assessment guidelines, transparency, and faculty development; governance of AI in assessment. |
| P225 | Educational Measurement as an AI-Native Profession | Panel on professional norms, training, accountability and standards of an AI-native measurement profession. |
| P238 | Trust Calibration in AI-Assisted Content Verification | How SMEs calibrate trust in AI verification concerns; professional review responsibilities. |
| P252 | Is AI a Foundational Competency or Does AI Elevate Aspects of Competencies? | Whether AI is a foundational competency for measurement professionals (NCME competencies). |
| P253 | Don’t Throw the BAIby Out with the Bathwater: Connecting FCEM to AI | Connecting NCME foundational competencies to AI; professional competencies. |
| P254 | AI Competencies in Education and Credentialing: Aligning Expectations with Evolv | AI competencies for measurement and credentialing professionals aligned with evolving practice. |
| P255 | Technical AI Competencies as Measurement Competencies | Technical AI competencies as measurement competencies; professional competency framework. |
| P310 | Responsible AI in Digital Assessment: Case Studies with the Duolingo English Tes | Responsible AI case studies at a testing program; governance of AI use in assessment practice. |
