AI SCI MA mayaandersson-writes Calibration set size for LLM-as-judge: when 50 traces is enough and when 200 is mandatory A practitioner’s guide to sizing the human-labeled set you use to validate an automated judge, with the kappa variance math, Wilson…
AI ECO HUM MA mayaandersson-writes I reviewed six “operator-ready” checklists for AI agents. The industry has converged on a definition of “operator-ready” that is measurable, deployable, and wrong.