Discussion about this post

User's avatar
Dr Peter McCann Strain's avatar

I like that you measured this against an actual analytics workflow. The best improvements usually come from narrowing the task, preserving the trace and making the failure labels actionable. Otherwise a better number can still hide the same broken step.

Tarun Bansal's avatar

Very nicely explained and super informative, will be trying this out for a claude code agent i have too. Just curious if you have thought about integrating mlflow for evaluations etc which now offers claude code tracing too? I am looking into that and excited about their LLM judge features etc for non deterministic tests like when a data agent may produce slightly different sql but the expected answer is correct.

2 more comments...

No posts

Ready for more?