<- Back to portfolio

Case study

RaschLab

A Python toolkit that reproduces a legacy item-response tool's output on real assessment datasets, row by row.

Role
Independent build, private repo
Stack
Python, Rasch / IRT, numpy, openpyxl

Problem and constraints

Item analysis for national assessment data ran through a legacy item-response binary: every run printed fixed-width text tables, and checking a new implementation against it meant reading the numbers by eye. The replacement had to consume the same control files and fixed-width response files the existing workflow already produced, and the reference tool's own convergence report showed its stop rules were not the ones a reimplementation would naturally pick.

Approach

Implemented dichotomous Rasch estimation in Python - person and item measures, fit statistics and the summary blocks - reading the control-file and fixed-width response-file formats the workflow already produced, and writing CSV plus an XLSX workbook (item table, person table, summary). Two estimation modes ship side by side: one mirrors the reference tool's path and stops at a calibrated threshold, the other is a Newton-Raphson implementation used as the reference-free check. Parity is measured rather than assumed: the vendor tool's own output for six anchored runs was parsed into golden fixtures and committed, so the comparison harness runs where the vendor export is absent, and every column of every table is gated instead of only the headline measure.

Measured result

Item-measure correlations exceed 0.9999 across all six runs; the worst per-run maximum difference in compat mode is 0.0314 logit, and the worst rank displacement is 3 positions. The stop threshold is calibrated against those runs rather than copied from the vendor manual, which is recorded in the repo as a deliberate deviation.

What Value Scope and source
item-measure correlation across the reference runs > 0.9999 six anchored runs, compared column by column against the reference tool's own tables
worst maximum difference per run, compat mode 0.0314 logit six runs covering 547 items in total
worst rank displacement 3 positions 2 positions or fewer on the other five runs
items covered by the golden fixtures 547 four instrument sections plus two revisions, 76 to 147 items each
reference data available without the vendor tool 6 runs aggregate per-iteration numbers only - no student data in the committed fixtures
runtime dependencies 2 numpy for estimation and fit statistics, openpyxl for the XLSX output

Stack

Python Rasch / IRT numpy openpyxl pytest