Top
Best
New

Posted by franze 10 hours ago

Show HN: Let your AI agents paint big arrows, boxes and text on your screen(github.com)
352 points | 150 comments
sicktriple 5 hours ago|
Man, just when it seemed like we had it all. Computers were cheap, efficient and powerful. Somehow we figured out a way to accomplish tasks we already had solved except now it's 1000x more expensive, requires the combined electricity of the entire world, is reliant on someone else's rented compute, and now I need a robot to tell me what button to press. What a time to be alive.
scotty79 24 minutes ago||
In 40 year software designers couldn't figure out that there should be a search function for every functionality that your computer provides.

When we got a form of such search, it searched all of the internet alongside with the functions of your software on your computer. You had to tell it manually what software you have and when it found something you had to fish out the description of what to click to get the functionality you are seeking and then click through some dumb ui to actually invoke it.

Agents are the first thing that can lead you directly from "I know what I want my computer to do." to "Actually doing it."

Success of agents is founded on the profound and sustained failure of all of the software designers and developers ever.

godelski 3 hours ago|||
For ages people have been saying "it's obvious" or "it's intuitive", without learning anything about design principles. Obvious to who?

People have been saying "we don't need docs", "nobody reads docs", "they'll figure it out". They experience the pain of a new employee onboarding but never treat this as a signal for the user experience. The user can't just walk over to the developer's desk. But luckily we got stack overflow, blogs, and Google. Somewhere someone explains it! The problem became finding it!

But hey, why do things the "hard" way when you can just throw money at the problem? Or better yet, dismiss the problem by calling the user an idiot.

You're right. What a time to be alive. We've created a world where this product is useful. And not even to because dark patterns exist, but simply because we've always wanted to avoid taking a step back and thinking or getting an outside view. Because it isn't obvious if a robot has to teach you what button to press. If it does, you probably should be ashamed (there are, as always, exceptions)

That said, I do think there's still a lot of utility you this. Those exceptions aren't uncommon. This'll help people learn programs like FreeCAD or Blender, or whatever. Where the complexity is naturally high. But also I do think we should recognize the silliness of many problems that this does solve that shouldn't be problems in the first place.

joquarky 53 minutes ago||
> Obvious to who?

Garage sale signs are a good example of this. The person making the sign knows what it says, so they can "read" it from farther away than someone who doesn't already know what it says.

alanbernstein 4 hours ago|||
Valid points, but this sounds extremely useful for learning complicated GUI applications like 3d modeling or media editing apps.
ASalazarMX 4 hours ago||
Valid point, but the solution is wasteful and overengineered. Imagine if videogames needed an external datacenter to run the tutorial levels.
arcanemachiner 2 hours ago|||
I am willing to bet you could run a model capable of using this on a 12GB 3060 + some RAM. (Qwen 3.6 35B)

Even the data centre using this probably uses less net energy, and costs less, to build one of these dumb arrows than the aggregate sum of the energy used to keep you alive while you Googled for the answer. (Including heating/cooling the building, powering the equipment growing the food you eat, etc.)

fasterik 3 hours ago|||
Sounds like an opportunity for a good engineer to come in and write something with the same functionality that runs locally and efficiently. The claim that it's wasteful only holds water if the non-wasteful solution exists and can accomplish the same tasks.
serf 4 hours ago|||
i've read enough 'shlemiel the painter' ports across hundreds of systems and thousands of examples that I must quickly and coldly dismiss the precept that computers were ever used efficiently.
Anon1096 3 hours ago|||
Now imagine what people back in the day thought about going from assembly to C. Or C to Java. Or, forgive me for even saying it, Java to Python.
sudo_cowsay 4 hours ago|||
On the other hand, (while you made a completely valid point,) some people who are new to a field, like young aspiring computer scientists who don't know how to do some stuff and need a guiding hand, will find this tremendously useful. Also, technology and compute power has been growing a tremendous amount over the past 5-10 years (now we have 20-100 billion transistors in chips!). <-- enough to sustain this wild use of compute

Indeed, what a time to be able.

CamperBob2 1 hour ago|||
If you look at the example animations, it's clear that Apple has brought this entirely upon themselves. Don't yell at some guys who are trying to fix it.
redanddead 3 minutes ago||
Human in the loop is the superior workflow!

/s

hn8726 9 hours ago||
I tried to read the "Does it need Screen Recording or Accessibility?" part, but it's slopped to the point I have no clue what it's trying to say. But if it can draw on top of permission prompts, what's stopping it from drawing box that hides the "decline" button and changing the "approve" button copy?
causal 8 hours ago||
I wonder if someday we will get to the point where Github repos are just markdown files describing the project and then you just let your own agent implement it because why the hell would I trust your agent's implementation?
jaggederest 4 hours ago|||
Here's a couple examples, I've seen others as well:

https://github.com/seb3773/ntfs-repair-rfc

https://github.com/Kotivskyi/screenshot-tool

I think it was more popular around the beginning of the year to mid-year

nater5000 7 hours ago||||
Seriously. I imagine this will be the case soon enough in some shape or form.

Whenever I see these vibecoded apps, my move is to just do what you've described: point an agent at it and tell it to reproduce it (usually with my own little customizations). It's a bit absurd, but I don't trust that they vibecoded the app as well as "I" could lol.

Abimelex 6 hours ago|||
A 100%! Just imagine you could describe every solution to a problem just using language! Of cause you would to make sure to avoid ANY missunderstanding, but therfore you could invent a language specified for avoiding ambiguity. I could imagine just to reduce the words to a very small corpus so everybody can remember it and have a very strict grammar so a program can effortlessly check correctness of its sentences.
ccozan 5 hours ago|||
I hope someone reads your post. I had to laugh hysterically when I reached the end. The level of sarcasm is too high.
joquarky 42 minutes ago|||
You have not experienced spec writing until you have written them in the original Klingon.
tempest_ 5 hours ago||||
This only works while tokens are artificially cheap. The gravy train could come to an end at some point.

The open models that are chasing the frontier labs will stop being open once things slow down and there is less incentive to undercut the front runners. Time will tell if GPU compute gets cheap enough to run stuff locally.

zaik 4 hours ago||||
This post makes me want to buy Nvidia stocks.
ccozan 5 hours ago|||
/goal build this but better
theropost 6 hours ago||||
Isn't it already kind of like that? Except the agent reads the code as if it's markdown. There's really no difference anymore, is there?
asdff 4 hours ago||||
With enough model drift that won’t even work over time. These files would have to be pinned to the intended model version and that is either used directly or emulated with a faithful emulator in a larger model.
cootsnuck 3 hours ago|||
I doubt that will be a problem. Models are different enough right now. If you can write a spec with enough detail and constraints today to get valid and comparable output from say GPT-6 Astra, Opus 5.5, DSV4F, Kimi K3, GLM 5.3, etc... Then I think there's a good chance that whatever SoTA coding LLMs everyone is using 3 years from now will also be able to implement that same spec.

Again, I think it heavily depends on if people are writing comprehensive specs with sufficient detail.

asdff 1 hour ago||
That depends on if you are really getting what you expect from the spec or you are actually relying on undefined behavior for your expected output from the spec. This is why we often pin software library versions: we may very well be relying on unintended or a bugged behavior to get out expected output.
rrr_oh_man 4 hours ago|||
llmpm
glitchc 5 hours ago||||
Sounds like Github is optional at that point. Why not just use Medium or WordPress?
asdff 4 hours ago|||
Just tell claude to share it with other claude users
lstodd 4 hours ago|||
reddit! use reddit!
m-s-y 5 hours ago||||
ding ding! we have a winner!
fennecfoxy 6 hours ago||||
Eh I think it'll be more like Gibson's defensive ICE in that ICE is "my swarm of agents scans your code for nasties, then compiles it from source".
glitchc 5 hours ago|||
Neither party owns their agents though. Is it simply then a matter of "my subscription is better than your subscription"?
pydry 6 hours ago|||
First you'd have to stop the agents from flagging hundreds of irrelevant nasties.

I think the next generation of exploit will be putting up code all over the internet that does x, y and z and then when asked to one shot whatever that code does the agents bake in the exploit that was indirectly included in their training corpus.

cyanydeez 8 hours ago|||
if all your models are in the cloud, why would you trust anything your agent builds
jerf 8 hours ago||
That sounds cynical today, but that's just because the AI models are currently outrunning enshittification. I'm already pondering personal plans about what to do when that turns around. We haven't seen enshittification yet that is going to be like the enshittification of AI. It may even deserve a new term of its very own, it's going to be such a big problem. The AI companies are leaving a lot of value on the table to entice us on to their systems but at some point that's going to turn around.
cyanydeez 7 hours ago||
The existential danger exists today: you're providing your entire business toolchain to models in the cloud, owned by businesses who are for profit entity. Even if the model itself doesn't care about your data, the business model does.

Everything your building is now streaming through a third party who could steal it, the same way Amazon steals open source software and repackages it.

All you need to do is look at palentir's forward deployed engineers. Once you let a third party observe your processed, you're cloneable and replaceable _as an entity_.

I think we can look at our own human psychology as it developed into social groups. If our brains were obviously deterministic, it would be easy for people to prey upon us via structured inputs. Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.

So these AI companies are definitely more dangerous than just the "AI is super intelligence". They're basically a intelligence platform everyone is wiring into.

That's today. Tomorrow, they'll be run by an MBA which is the enshittification.

mat_b 1 hour ago|||
> Some level of randomness in our thought process is almost evolutionarily required to avoid becoming part of the larger super structure, similar to an ant colony.

Not really. Ants are superorganisms, they share more than half of their DNA with each other. You could think of each ant as being like a cell in a human.

jack_pp 6 hours ago|||
software businesses are not like watches (recently saw a youtube video about fakes being made virtually identical at 10% of the price).

If you can make a facebook clone it doesn't mean you can steal facebook's lunch. there are exceptions of course.

hannasanarion 7 hours ago|||
The fact that the utility doesn't give it the ability to do that?

It makes sense to not trust AI models with the ability to read and alter your screen for many many many reasons, but the developer of the tool knows that as well. A tool that lets your ai do something doesn't also let it do anything.

So, as always, it's a question of if you trust the software producer. You're giving them the ability to draw on your screen too. Would you have the same fear if there were no AI involved?

For stuff like this I imagine AI misbehavior as a novel failure trigger not a novel failure type. It can only do the harm that the software was able to do on its own already.

SwtCyber 7 hours ago|||
Nothing stops it, except the fact that the agent is already executing arditrary code in your shell. If it's malicious, it'll just steal your shh keys directly instead of bothering with button masking
smugglerFlynn 6 hours ago|||
Don't tell me you are reading readmes with your eyes in 2026. Blasphemy!
tkdb 8 hours ago||
...and there we have it. Slop assumes a verb form.
andai 8 hours ago|||
Ensloppification!
layer8 8 hours ago|||
“Slop” has been a verb for a long time: https://en.wiktionary.org/wiki/slop#Verb
tkdb 7 hours ago|||
You know the rules. Now that it's AI slopping it's new and innovative.
tottenhm 5 hours ago||||
Some words reverberate.
trollbridge 7 hours ago|||
Indeed, I’ve heard it on the context of feeding animals my entire life.
internet101010 1 hour ago||
The worst trend in UX in the last decade is the endless "Got it!" popups and feature notifications that distract the user from what they were trying to do.

I don't know why anyone would ever willingly want this.

tangotaylor 7 hours ago||
"It is an arrow, so we spent an unreasonable amount of time on how it looks."

Brilliant. This is exactly the kind of content I seek when I visit Hacker News.

Truly art.

socializer 6 hours ago|
Have you actually looked at this? In the first screenshot, three of the five arrows are obviously and badly misaligned, have incorrect labels, and it's just a big tangled mess (edit: the author stealth-replaced the screenshot, but the original is here: https://raw.githubusercontent.com/franzenzenhofer/big-arrow-...). The LLM-generated text may be saying one thing, but there's clearly zero effort spent on... anything. There's no punchline, there's no aesthetic angle, there's no conceivable purpose.

It's just the epitome of living in an era where code is so cheap to generate that we don't invest any effort into even fully fleshing out the idea. It's just "Claude, make me a program that draws arrows on screen". And HN, I guess, is dominated by two groups: one still stuck in the mentality of "code = effort" and trying to recognize that; and another stuck in the mentality of "AI = cool".

usrbinbash 8 hours ago||
SO the point of this is ... what exactly?

A big arrow to an interface element which ... has a label that explains what it does?

So...a label for a label?

hannasanarion 6 hours ago||
It seems silly but I think there's a good use case for this:

Helping people deal with bad UX.

Everybody who has worked in a traditional corporate environment has encountered poorly tutorialized, poorly labeled, counterintuitive user interfaces where the only way to know how to get something done is to have somebody sit over your shoulder and tell you which buttons to click.

Oh, you want to figure out which OIM entitlements your coworker has that you don't to debug an access issue? You just need to know that you have to click the "compliance" button to get there. Why is it called "compliance" when what it actually does is let you export comparisons from the user tables? Because fuck you that's why. Memorize it.

An AI assistant that can read the procedures for you, or sometimes even read the source code, or the website code, can show you how to accomplish something that the UI on its own fails to do.

wartywhoa23 3 hours ago||
Wait, I was under impression that UX was to be sorted out by AI just like the code is?
voidUpdate 8 hours ago|||
It's so your agent can make a big arrow on the screen saying "click this, human" so that it can keep doing things
andyfilms1 8 hours ago||
I love living in the future!
mistersquid 7 hours ago|||
I ran two experiments which require the `claude` CLI tool be installed.

For the first experiment, Claude stepped me through how to edit its settings.json to remove warnings from its configuration: a little clunky as I had to respond to Claude's CLI prompts to get it to auto run.

In the second experiment, Claude walked me through how to use Apple's Compressor.app to speed up a video. It walked me through dragging the file into compressor and clicking on the appropriate UI elements to do this.

So, one use case for the bigarrow skill enables Claude to step users through a relatively complex UI to accomplish a task.

There are likely more efficient ways to accomplish the tasks at hand (e.g. ffmpeg for the second experiment), but this use case is interesting if using a GUI tool is required.

cobbal 5 hours ago|||
Finally we've built the reverse centaur from Cory Doctorow's classic novel "Don't Become the Reverse Centaur"
ghm2180 8 hours ago||
Glad you asked. Pointing my aging aunt to the right place on the screen to click without having to take control of her laptop, of course.
arshxyz 10 hours ago||
The README is geared towards technical people (complete with the HN screenshot) but when I see a tool like this all I can think of is how helpful this would be for my mom when I'm trying to tell her how to download and print a document over the phone
conception 9 hours ago||
Honestly the giant arrow annotation Zoom has makes it worth any amount of money compared to the competition.
prmoustache 3 hours ago||
Except your mom just have to tell her AI agent to download and print that document.

She doesn't have to call you anymore.

kogus 4 hours ago||
My first reaction was similar to the reaction I'd have if you told me that cockroaches had learned to unlock doors and stand on their hind legs. But then I thought about accessibility, and the ways this could be used to help technologically illiterate or disabled people, and I thought again.

Do you remember when PCs used to come with a completely soup-to-nuts tutorial that would talk to you like you had never seen a PC before? Things like this: https://www.youtube.com/watch?v=3ScS4OYDfHE

This kind of baked-in interactivity could really help in a training or disability context.

isoprophlex 10 hours ago||
Literally unusable as it is. Some minimal extra features this would need:

- rainbow dripping arrows

- angrily pointing arrows

- flame-surrounded text boxes with particle effects

- the ability for the agent to play airhorn.wav at max volume, overriding existing volume or mute settings

EDIT: ayy lmao https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...

htrp 9 hours ago||
This is where the agent says did not understand sarcasm coded and shipped features
DonHopkins 9 hours ago||
Last time I accidentally said something sarcasticly over-ambitious to an LLM, it shipped this popup callout tooltip feature on a PDP-7 Type 340 vector graphics display emulator that shows you the meaning of the drawing you're pointing at, as well as the address of the instruction that drew it.

https://hyperties.org/cabinet/symelec/

PIXIE and FORTH on PDP-7 Cabinet Emulator with Type 340 Vector Graphics Display:

https://www.youtube.com/watch?v=lo8kdY-5i6c

dihinbutt 6 hours ago|||
W developer
thih9 10 hours ago||
At the risk of stating the obvious - let's not do that, the goal of this repo is to be useful and not to give agents the power of the `<blink>` tag.
voidUpdate 8 hours ago|||
> ""I need you, and you're making coffee." --say reads the sign aloud. Your Mac will literally call you back to your desk."

It's already got the power to be obnoxious at you

koalacola 9 hours ago||||
Oh dear, they were making a joke.
yen223 9 hours ago||
if only there was a way to make a subtle thing obvious
isoprophlex 9 hours ago||
such as... angry flaming rainbow textboxes and arrows?
ale42 8 hours ago||
I thought that the dripping rainbow ones were enough. Maybe you have to ask for rainbow unicorns flying on the screen.
jaapz 8 hours ago||||
the repo is one big joke, of course they should add this
thih9 6 hours ago||
[dead]
DonHopkins 9 hours ago||||
For the humor impared, it would also be useful to have a colorful animated "WHOOSH" overlay with sound effects for every time a deadpan joke goes over your head. ;)

Maybe isoprophlex will add that to his PR!

ipsod 9 hours ago|||
under_construction.gif
isoprophlex 9 hours ago||
just submitted the airhorn PR; ~second rainbow arrow slop grenade incoming~ BOOM slop cannon fired
franze 3 hours ago|||
thx, not merged but great PR https://github.com/franzenzenhofer/big-arrow-on-the-screen/p...
ipsod 26 minutes ago||
Lol, that's awesome.
thih9 6 hours ago|||
[dead]
lbreakjai 9 hours ago||
I would pay good money for something like this on iPad. It wouldn't even need to be agent-driven, just a big "I want to make a bank transfer" button, that would launch the correct app and guide through the interface.

That would be a godsent for those of us with aging parents.

ghm2180 9 hours ago|
> That would be a godsent for those of us with aging parents.

Right yeah? I mean I have freagin remote desktop installed on all my parent's devices. I am the agent for them every time they call frantically for help.

It's Time to redesign the apple genius bar for the modern aging boomer using this.

alansaber 5 hours ago|
OP has inspired me to write up my tedious thoughts on agent GUIs if of interest https://news.ycombinator.com/item?id=50022688
More comments...