The Counterfeit Test for AI Metrics
We had an AI quality metric that was stable across model tiers, improved when we refined the system, and measured something tied to proprietary knowledge. Then we built a counterfeit with almost none of the relevant expertise and it scored 0.99. A metric is not trustworthy because the real system passes. It becomes trustworthy when the fake fails.