Jagged AGI: o3, Gemini 2.5, and everything after

(www.oneusefulthing.org)

Show context

mellosouls ◴[20 Apr 25 17:44 UTC] No.43745240[source]▶

The capabilities of AI post gpt3 have become extraordinary and clearly in many cases superhuman.

However (as the article admits) there is still no general agreement of what AGI is, or how we (or even if we can) get there from here.

What there is is a growing and often naïve excitement that anticipates it as coming into view, and unfortunately that will be accompanied by the hype-merchants desperate to be first to "call it".

This article seems reasonable in some ways but unfortunately falls into the latter category with its title and sloganeering.

"AGI" in the title of any article should be seen as a cautionary flag. On HN - if anywhere - we need to be on the alert for this.

replies(13): >>43745398 #>>43745959 #>>43746159 #>>43746204 #>>43746319 #>>43746355 #>>43746427 #>>43746447 #>>43746522 #>>43746657 #>>43746801 #>>43749837 #>>43795216 #

daxfohl ◴[20 Apr 25 21:30 UTC] No.43746657[source]▶

>>43745240 #

Until you can boot one up, give it access to a VM video and audio feeds and keyboard and mouse interfaces, give it an email and chat account, tell it where the company onboarding docs are and expect them to be a productive team member, they're not AGI. So long as we need special protocols like MCP and A2A, rather than expecting them to figure out how to collaborate like a human, they're not AGI.

The first step, my guess, is going to be the ability to work through github issues like a human, identifying which issues have high value, asking clarifying questions, proposing reasonable alternatives, knowing when to open a PR, responding to code review, merging or abandoning when appropriate. But we're not even very close to that yet. There's some of it, but from what I've seen most instances where this has been successful are low level things like removing old feature flags.

replies(3): >>43746758 #>>43747095 #>>43747467 #

rafaelmn ◴[20 Apr 25 21:50 UTC] No.43746758[source]▶

>>43746657 #

Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models. Having a specialized interface protocol is orthogonal to capability. Just like you could theoretically write code in a proportional font with notepad and run your tools through windows CMD - having an editor with syntax highlighting and monospaced font helps you read/navigate/edit, having tools/navigation/autocomplete etc. optimized for your flow makes you more productive and expands your capability, etc.

If I forced you to use unnatural interfaces it would severely limit your capabilities as well because you'd have to dedicate more effort towards handling basic editing tasks. As someone who recently swapped to a split 36key keyboard with a new layout I can say this becomes immediately obvious when you try something like this. You take your typing/editing skills for granted - try switching your setup and see how your productivity/problem solving ability tanks in practice.

replies(3): >>43747058 #>>43747819 #>>43752611 #

1. esperent ◴[21 Apr 25 01:28 UTC] No.43747819[source]▶

>>43746758 #

> Just because we rely on vision to interface with computer software doesn't mean it's optimal for AI models

This is true but AGI means "Artificial General Intelligence". Perhaps it would be even more efficient with certain interfaces, but to be general it would have to at least work with the same ones as humans.

Here's some things that I think a true AGI would need to be able to do:

* Control a general purpose robot and use vision to do housework, gardening etc.

* Be able to drive a car - equivalent interfaces to humans might be service motor controlled inputs.

* Use standard computer inputs to do standard computer tasks

And this list could easily be extended.

If we have to be very specific in the choice of interfaces and tasks that we give it, it's not a general AI.

At the same time, we have to be careful at moving the goalposts too much. But current AI are limited to what can be returned in a small number of interfaces (prompt with text/image/video & return text/image/video data). This is amazing, they can sound very intelligent while doing so. But it's important not to lose sight of what they still can't do well which is basically everything else.

Outside of this area, when you do hear of an AI doing something well (self driving, for example) it's usually a separate specialized model rather than a contribution towards AGI.

replies(2): >>43747924 #>>43753643 #

2. mNovak ◴[21 Apr 25 01:59 UTC] No.43747924[source]▶

>>43747819 (TP) #

By this logic disabled people would not class as "Generally Intelligent" because they might have physical "interface" limitations.

Similarly I wouldn't be "Generally Intelligent" by this definition if you sat me at a Cyrillic or Chinese keyboard. For this reason, I see human-centric interface arguments as a red herring.

I think a better candidate definition might be about learning and adapting to new environments (learning from mistakes and predicting outcomes), assuming reasonable interface aids.

replies(2): >>43748508 #>>43748532 #

3. esperent ◴[21 Apr 25 04:30 UTC] No.43748508[source]▶

>>43747924 #

> Similarly I wouldn't be "Generally Intelligent" by this definition if you sat me at a Cyrillic or Chinese keyboard

Would you be able to be taught to use those keyboards? Then you're generally intelligent. If you could not learn, then maybe you're not generally intelligent?

Regarding disabled people, this is an interesting point. Assuming that we're talking about physical disabilities only, disabled people are capable of learning how to use any standard human inputs. It's just the physical controls that are problematic.

For an AI, the physical input is not the problem. We can just put servo motors on the car controls (steering wheel, brakes, gas) and give it a camera feed from the car. Given those inputs, can the AI learn to control the car as a generally intelligent person could, given the ability to use the same controls?

4. vczf ◴[21 Apr 25 04:35 UTC] No.43748532[source]▶

>>43747924 #

If all we needed was general intelligence, we would be hiring octopuses. Human skills, like fluency in specific languages, are implicit in our concept of AGI.

replies(1): >>43750281 #

5. ◴[21 Apr 25 10:30 UTC] No.43750281{3}[source]▶

>>43748532 #

6. ctoth ◴[21 Apr 25 16:16 UTC] No.43753643[source]▶

>>43747819 (TP) #

So I am a blind human. I cannot drive a car or use a camera/robot to do housework (I need my hands to see!) Am I not a general intelligence?

replies(1): >>43757739 #

7. esperent ◴[21 Apr 25 23:55 UTC] No.43757739[source]▶

>>43753643 #

I replied this to another comment, but I'll put it here: your limitation is physical. You have standard human intelligence, but you're lacking a certain physical input (vision). As a generally intelligent being, you will compensate for the lack of vision by using other senses.

That's different to AIs, which we can hook up to all kinds of inputs: cameras, radar, lidar, car controls, etc. For the AI the lack of input is not the limitation. It's whether they can do anything with an arbitrary input/control, like a servo motor controlling a steering wheel, for example.

To look at it another way, if an AI can operate a robot body by vision, then we suddenly removed the vision input and replaced it with a sense of touch and hearing, would the AI be able to compensate? If it's an AGI, then it should be able to. A human can.

On the other hand, I wonder if we humans are really as "generally intelligent" as we like to think. Humans struggle to learn new languages as adults, for example (something I can personally attest to, having moved to Asia as an adult). So, really, are human beings a good standard by which to judge an AI as AGI?

↑