---
id: "1954224651443544436"
requested_id: "1954224651443544436"
status: ok
level: 1
user: "karpathy"
name: "Andrej Karpathy"
created_at: "2025-08-09T16:53:59.000Z"
in_reply_to: null
quoted: null
source: tweet-result
url: "https://x.com/karpathy/status/1954224651443544436"
via:
  - quote: "1954242913875140960"
---

I'm noticing that due to (I think?) a lot of benchmarkmaxxing on long horizon tasks, LLMs are becoming a little too agentic by default, a little beyond my average use case.

For example in coding, the models now tend to reason for a fairly long time, they have an inclination to
