Problem and constraints
Item analysis for national assessment data ran through a legacy item-response binary: every run printed fixed-width text tables, and checking a new implementation against it meant reading the numbers by eye. The replacement had to consume the same control files and fixed-width response files the existing workflow already produced, and the reference tool's own convergence report showed its stop rules were not the ones a reimplementation would naturally pick.
Approach
Implemented dichotomous Rasch estimation in Python - person and item measures, fit statistics and the summary blocks - reading the control-file and fixed-width response-file formats the workflow already produced, and writing CSV plus an XLSX workbook (item table, person table, summary). Two estimation modes ship side by side: one mirrors the reference tool's path and stops at a calibrated threshold, the other is a Newton-Raphson implementation used as the reference-free check. Parity is measured rather than assumed: the vendor tool's own output for six anchored runs was parsed into golden fixtures and committed, so the comparison harness runs where the vendor export is absent, and every column of every table is gated instead of only the headline measure.
Measured result
Item-measure correlations exceed 0.9999 across all six runs; the worst per-run maximum difference in compat mode is 0.0314 logit, and the worst rank displacement is 3 positions. The stop threshold is calibrated against those runs rather than copied from the vendor manual, which is recorded in the repo as a deliberate deviation.
| What | Value | Scope and source |
|---|---|---|
| item-measure correlation across the reference runs | > 0.9999 | six anchored runs, compared column by column against the reference tool's own tables |
| worst maximum difference per run, compat mode | 0.0314 logit | six runs covering 547 items in total |
| worst rank displacement | 3 positions | 2 positions or fewer on the other five runs |
| items covered by the golden fixtures | 547 | four instrument sections plus two revisions, 76 to 147 items each |
| reference data available without the vendor tool | 6 runs | aggregate per-iteration numbers only - no student data in the committed fixtures |
| runtime dependencies | 2 | numpy for estimation and fit statistics, openpyxl for the XLSX output |