←back to thread

AI agent benchmarks are broken

(ddkang.substack.com)
181 points neehao | 1 comments | | HN request time: 0.426s | source
1. let_tim_cook_ ◴[] No.44532633[source]
Are any authors here? Have you looked at AppWorld? https://appworld.dev