# Scoring the supplied fixtures

From this resources directory run `python3 score_results.py --output evaluation_report.json`. Python 3.10+; standard library only. The script checks case/result IDs and computes declared-label rates. Inspect semantic labels against the sources yourself. This is not a live model benchmark.
