8 ms·
Tests are called "evals" (evaluations) in the AI product development world. Basically you let humans review LLM output or feed it to another LLM with instructio
by sk7 9mo ago
Tests are called "evals" (evaluations) in the AI product development world. Basically you let humans review LLM output or feed it to another LLM with instructions how to evaluate it.
https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-complete https://www.lennysnewsletter.com/p/beyond-vibe-checks-a-pms-...
- azemetre 9mo agoInteresting, never really thought of it outside of this comment chain but I'm guessing approaches like this hurt the typical automated testing devs would do but seeing how this is MSFT (who already stopped having dedicated testing roles for a good while now, rip SDET roles) I can only imagine the quality culture is even worse for "AI" teams.
- ethbr1 9mo agoYes. Because why would there ever be a problem with a devqaops team objectively assessing their own work's effectiveness?