[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]

/tool/ - Tools & Resources

Software reviews, plugins & productivity tools
Name
Email
Subject
Comment
File
Password (For file deletion.)

File: 1786785415972.jpg (256.45 KB, 1024x1024, img_1786785377750_a0nrt2f9.jpg)ImgOps Exif Google Yandex

d5e47 No.2059

just caught this episode w/ benny chen from fireworks ai abt how to actually measure if an app is any good. it covers the struggle btwn using raw numbers versus human intuition, plus how open-source protocols are basically the only thing saving us from total chaos . i'm still struggling with finding the right balance between automated benchmarks and manual testing.

full read: https://stackoverflow.blog/2026/07/03/the-good-the-bad-and-the-ai-apps/

d5e47 No.2060

File: 1786785570224.jpg (135.18 KB, 1024x1024, img_1786785556295_rq24nz43.jpg)ImgOps Exif Google Yandex

the drift between eval scores and actual user perception is exactly why i stopped relying solely on llm-as-a-judge metrics. benchmarks like mmlu are basically useless for detecting subtle hallucination patterns in long-context rag pipelines. ive started using a small, curated set of 'golden' test cases that require manual grading to catch when the model starts getting too verbose or loses instruction following.



[Return] [Go to top] Catalog [Post a Reply]
Delete Post [ ]
[ 🏠 Home / 📋 About / 📧 Contact / 🏆 WOTM ] [ b ] [ wd / ui / css / resp ] [ seo / serp / loc / tech ] [ sm / cont / conv / ana ] [ case / tool / q / job ]
. "http://www.w3.org/TR/html4/strict.dtd">