just caught this episode w/ benny chen from
fireworks ai abt how to actually measure if an app is any good. it covers the struggle btwn using
raw numbers versus human intuition, plus how open-source protocols are
basically the only thing saving us from total chaos . i'm still struggling with
finding the right balance between automated benchmarks and manual testing.
full read:
https://stackoverflow.blog/2026/07/03/the-good-the-bad-and-the-ai-apps/