Posted by franze 10 hours ago
When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.
Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."
Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.
People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!
But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.
You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)
That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.
Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.
Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)
Indeed, what a time to be able.
/s
https://github.com/seb3773/ntfs-repair-rfc
https://github.com/Kotivskyi/screenshot-tool
I think it was more popular around the beginning of the year to mid-year
Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.
The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.
Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.
I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.
Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.
All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.
I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.
So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.
That's today. Tomorrow, they'll be run by an MBA which is the enshittification.
Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.
If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.
It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.
So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?
For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.
I don't know why anyone would ever willingly want this.
Brilliant. This is exactly the kind of content I seek when I visit Hacker News.
Truly art.
It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".
A big arrow to an interface element which ... has a label that explains what it does?
So...a label for a label?
Helping people deal with bad UX.
Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.
Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.
An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.
For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.
In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.
So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.
There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.
She doesn't have to call you anymore.
Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE
This kind of baked-in interactivity could really help in a training or disability context.
- rainbow dripping arrows
- angrily pointing arrows
- flame-surrounded text boxes with particle effects
- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings
EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
https://hyperties.org/cabinet/symelec/
PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:
It's already got the power to be obnoxious at you
Maybe isoprophlex will add that to his PR!
That would be a godsent for those of us with aging parents.
Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.
It's Time to redesign the apple genius bar for the modern aging boomer using this.