Benchmarks: The Good, the Bad, and the Ugly — Ali Khial, G2i
AI benchmark prompts are so unrealistic that experienced engineers would never write them in practice
“No one writes prompts like these ever.”
G2i's Director of AI and ML walked through his team's investigation of AI benchmark design, finding that benchmark task prompts are so artificial that real engineers confirmed they would never write them. He simplifies the benchmark pipeline to its fundamentals — prompt, model/agent, verifier/grader, harness — arguing that benchmark quality depends on all components working correctly. The content is cut off before conclusions, limiting the signal value.