Posted by franze 12 hours ago
(OP, nice project! Sorry for my rant.)
But it's a subpar experience compared to just opening github on the browser.
Why does this need to be a skill? Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows” and you get the same thing without burning tens of thousands of tokens in context.
> burning tens of thousands of tokens in context.
~1400?
> Literally just ask Claude to “call out actions in these screenshots with numbered stylized arrows”
Is there a macos native thing it uses? Or are you talking about screenshots? This draws on your screen instead.
Does one need 4 programming languages to draw something on a mac?
> "It's this window, not that one." You have 14 Chrome windows. The agent knows which one it means: --window "Google Chrome:Pull request". It even picks the right tab: --app "Google Chrome:Pull request".
> Remote help. "No, the other gear icon." Point at it instead of describing it.
To me this feels like it takes away from what the human is supposed to do (read, understand the consequences of the action, then.. consent or abort)
There is a reason your AI Agent won't automate these clicks for you