tech
Benchmarks Don't Know Your Job
Plus: the number KateBench taught us to count, and a six-agent solar crew helping decide whether to run the dryer

TL;DR
- Companies lack clear methods to measure AI's impact on employee time savings and work trustworthiness.
- Public benchmarks and spending on models do not equate to real-world effectiveness.
- Organizations need to create tests tailored to their specific work processes.
- Examples of AI evaluation include cloning an editor and using agent crews for practical decision-making.