Jagged AGI: o3, Gemini 2.5, and everything after

(www.oneusefulthing.org)

265 points ctoth | 1 comments | 20 Apr 25 14:55 UTC | HN request time: 0.25s | source

Show context

mellosouls ◴[20 Apr 25 17:44 UTC] No.43745240[source]▶

The capabilities of AI post gpt3 have become extraordinary and clearly in many cases superhuman.

However (as the article admits) there is still no general agreement of what AGI is, or how we (or even if we can) get there from here.

What there is is a growing and often naïve excitement that anticipates it as coming into view, and unfortunately that will be accompanied by the hype-merchants desperate to be first to "call it".

This article seems reasonable in some ways but unfortunately falls into the latter category with its title and sloganeering.

"AGI" in the title of any article should be seen as a cautionary flag. On HN - if anywhere - we need to be on the alert for this.

replies(13): >>43745398 #>>43745959 #>>43746159 #>>43746204 #>>43746319 #>>43746355 #>>43746427 #>>43746447 #>>43746522 #>>43746657 #>>43746801 #>>43749837 #>>43795216 #

daxfohl ◴[20 Apr 25 21:30 UTC] No.43746657[source]▶

>>43745240 #

Until you can boot one up, give it access to a VM video and audio feeds and keyboard and mouse interfaces, give it an email and chat account, tell it where the company onboarding docs are and expect them to be a productive team member, they're not AGI. So long as we need special protocols like MCP and A2A, rather than expecting them to figure out how to collaborate like a human, they're not AGI.

The first step, my guess, is going to be the ability to work through github issues like a human, identifying which issues have high value, asking clarifying questions, proposing reasonable alternatives, knowing when to open a PR, responding to code review, merging or abandoning when appropriate. But we're not even very close to that yet. There's some of it, but from what I've seen most instances where this has been successful are low level things like removing old feature flags.

replies(3): >>43746758 #>>43747095 #>>43747467 #

rafaelmn ◴[20 Apr 25 21:50 UTC] No.43746758[source]▶

>>43746657 #

Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models. Having a specialized interface protocol is orthogonal to capability. Just like you could theoretically write code in a proportional font with notepad and run your tools through windows CMD - having an editor with syntax highlighting and monospaced font helps you read/navigate/edit, having tools/navigation/autocomplete etc. optimized for your flow makes you more productive and expands your capability, etc.

If I forced you to use unnatural interfaces it would severely limit your capabilities as well because you'd have to dedicate more effort towards handling basic editing tasks. As someone who recently swapped to a split 36key keyboard with a new layout I can say this becomes immediately obvious when you try something like this. You take your typing/editing skills for granted - try switching your setup and see how your productivity/problem solving ability tanks in practice.

replies(3): >>43747058 #>>43747819 #>>43752611 #

raducu ◴[21 Apr 25 14:47 UTC] No.43752611[source]▶

>>43746758 #

> Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models.

It's optimal for beings that have general purpose inteligence.

> would severely limit your capabilities as well because you'd have to dedicate more effort towards handling basic editing tasks

Yes, but humans will eventually get used to it and internalize the keyboard, the domain language, idioms and so on and their context gets pushed to long term knowledge overnight and thei short term context gets cleaned up and they get bettet and better at the job, day by day. AI starts very strong but stays at that level forever.

When faced with a really hard problem, day after day the human will remember what he tried yesterday and parts of that problem will become easier and easier for the human, not so for the AI, if it can't solve a problem today, running it for days and days produces diminishing returns.

That's the General part of human intelligence -- over time it can aquire new skills it did not have yesterday, LLMs can't do that -- there is no byproduct of them getting better/aquiring new skills as a result of their practicing a problem.

replies(2): >>43753302 #>>43753617 #

ctoth ◴[21 Apr 25 16:14 UTC] No.43753617[source]▶

>>43752611 #

> It's optimal for beings that have general purpose inteligence [Sic].

Hi. I'm blind. I would like to think I have general-purpose intelligence thanks.

And I can state that interfacing with vision would, in fact, be suboptimal for me. The visual cortex is literally unformed. Yet somehow I can perform symbolic manipulations. Converse with people. Write code. Get frustrated with strangers on the Internet. Perhaps there are other "optimal" ways that "intelligent" systems can use to interface with computers? I don't know, maybe the accessibility APIs we have built? Maybe MCP? Maybe any number of things? Data structures specifically optimized for the purpose and exchanged directly between vastly-more-complex intelligences than ourselves? Do you really think that clicking buttons through a GUI is the one true optimal way to use a computer?

replies(2): >>43753763 #>>43754257 #

1. daxfohl ◴[21 Apr 25 17:20 UTC] No.43754257[source]▶

>>43753617 #

Of course not. The visual part is window dressing on the argument. The real point is, before declaring AGI, I think the way we interact with these agents needs to be more like human to human interaction. Right now, agents generally accept a command, figure out which from a small number of MCPs that have been precoded for it to use, do that thing you wanted right or wrong, the end. If it does the right thing, huge confirmation bias that it's AGI. Maybe the MCP did most of the real work. If it doesn't, well, blame the prompt or maybe blame the MCPs are lacking good descriptions or something.

To get a solid read on AGI, we need to be grading them in comparison to a remote coworker. That they necessarily see a GUI is not required. But what is required is that they have access to all the things a human would, and don't require any special tools that limit their search space to a level below what a human coworker would have. If it's possible for a human coworker to do their whole job via console access, sure, that's fine too. I only say GUI because I think it'd actually be the easiest option, and fairly straightforward for these agents. Image processing is largely solved, whereas figuring out how to do everything your job requires via console is likely a mess.

And like I said, "using the computer", whether via GUI or screen reader or whatever else, isn't going to be the hard part. The hard part is, now that they have this very abstract capability and astronomically larger search space, it changes the way we interact with them. We send them email. We ping them on Slack. We don't build special baby mittens MCPs and such for them and they have to enter the human world and prove that they can handle it as a human would. Then I would say we're getting closer to AGI. But as long as we're building special tools and limiting their search space to that limited scope, to me it feels like we're still a long way off.

↑