Top
Best
New

Posted by franze 11 hours ago

Show HN: Let your AI agents paint big arrows, boxes and text on your screen(github.com)
352 points | 150 commentspage 2
cyberjunkie 10 hours ago|
I'm just as impressed by this as any other LLM-generated project.
priyashunt 6 hours ago||
Very VERY USABLE FOR old PEOPLE. I would pay good money for something like this on iPad!
vessenes 11 hours ago||
Interesting. When I read the headline I imagined this would be a sort of thinking trace booster -- letting the agent focus its own attention on different parts of the screen. But this is cool in a different way. I bet agentic harnesses would find it useful for communicating with other agents / themselves as well.
alansaber 10 hours ago||
This might be goofy, but it underscores that there's potential for more visual agent UIUX than reading off a sidebar/opening modals.
melvinroest 10 hours ago|
[dead]
melvinroest 10 hours ago||
My message to the world is that LLMs should be able to point anything they see in the application they're in or even the whole computer (if you give it that kind of access).

For web apps, WebMCP is a way where you can make a frontend way more discoverable to an LLM rather than that it is going to read all your code. I sometimes create this in my apps at home and sometimes in my apps at work and it makes LLM assistants way more usable. It also helps if they can highlight certain things of a visual.

We humans can do this too on paper. We can point at something, we can highlight something with a marker. Why not give an LLM these capabilities as well?

FinnLobsien 11 hours ago||
This could be great for documentation. Screenshots in docs are frequently useless because they show me a screen and say "click X" where I still have to search X visually. And I could just to dhat in the other tab I have open.
peaxkl 9 hours ago|
It doesn’t make sense that you have to read a whole article and then still search for the buttons in the UI afterwards. And with longer articles, you always have to keep the article open next to your product to follow the whole flow.

We built something to help with that [1].

[1] https://www.happysupport.ai/en/in-app-messaging

alexpotato 8 hours ago||
> Arrows have existed since roughly the Paleolithic.

Whenever I use web tools that don't have "deep linking" [0], I love to throw out this quote:

"Have we thought about using hyperlinks? They are these SUPER useful things that were invented in Switzerland back in the 1990s."

0 - https://en.wikipedia.org/wiki/Deep_linking

swframe2 6 hours ago||
I want this for the visualization of very difficult to solve application evolution. This is for problems which have no known solution and are too complicated for an agent to figure out on its own. My current problem: can alphafold and related tools figure out the function of a gene that so far is unknown. Think of the monte carlo tree search in the alphago explanation videos. What if there was visualization that showed how the policy and value models worked so you can spot their flaws. I want the agent to build an attempt at a solution, then build a visualization of it, then run the solution and show me what is it up to. I want to it pause and explain its state so I can see exactly where and why the solution fails.
satyanash 11 hours ago||
Am I missing something here?

What terminal / agentic workflows spawn the demonstrated dialogue boxes that require the User to click/reject the action? Aren't most such flows actually inline UIs? And finally, if the core issue is that these confirmation boxes are tied to the terminal that triggered them, which does not autofocus, what's the point of said "big arrow" that is also lurking behind without focus?

If the said "big arrow" automatically gains focus, isn't the real fix here to just make the dialogue boxes themselves gain focus automatically? Both are similarly disruptive anyway.

inanutshellus 11 hours ago||
The first example (of HN) is the one that feels like it has the most potential to me.

"Teach me to do this" kinda stuff. "Guiding agent" rather than "doing agent".

Honestly, @franze, if you're reading this, maybe update your screenshots to show an agent in tutorial mode on some complicated app?

ravila4 9 hours ago||
I think this would’ve been very handy to me a couple years ago when I was learning to use Blender and asking LLMs for help performing certain actions like “how do I display the normals of all the vertices in my mesh?” I spent a lot of time trying to figure out which button the model was talking about.
dr_kiszonka 6 hours ago|
I take this opportunity to shame GitHub for completely ignoring the mobile experience in their own Android app. The project's README is pages upon pages of largely blank space. In general, the app does not render mermaid diagrams and does not allow for zooming in, so smaller pictures are unusable. Even if you access pictures directly in a repo's source, you still can't zoom in. iPython Notebooks, which are extremely common in data science, are an "unsupported file type" and are not rendered.

(OP, nice project! Sorry for my rant.)

bel8 6 hours ago|
I'm forced to use GH mobile app because of 2FA.

But it's a subpar experience compared to just opening github on the browser.

githubnuoo 5 hours ago||
it was driving me mad too. you can disable github.com opening in the github app: app-info/set-as-default/open-supported-links
More comments...