Posted by blazarquasar 1 day ago
> We want to make dealing with agentic infrastructure easier so you can focus on your work. AX is designed with an uncompromising focus on ergonomics, rapid iteration, and joyful workflows for both application developers and AI researchers.
On the the other hand, the readme quickstart section says
> You need a Kubernetes cluster, ko (brew install ko), a container registry your cluster can pull from, and a reachable Agent Substrate Control API (in-cluster default: api.ate-system.svc.cluster.local:443).
Call me old-fashioned but I don't find this "easier". Maybe it's easier in the same way that Kubernetes itself is easier than managing VMs and container deployments at massive scale without such a tool. But there's a vast chasm between what this tool is being sold as and what it actually is.
I don't really like the oversubscription of agent pods though, as you can no longer trust the k8s pod identity as being from a singular workload. Haven't seen a solution to this for ax yet and it is a barrier to adoption for us.
(This is work in flight, but it will land within a few weeks)
I believe the substrate egress-gateway needs to handover the internal SPIFFE one to an external system (e.g. Entra Agent ID). Not sure if that should be part of substrate or kagent/ax/..
lmao found the dev who's only ever worked on the dev side of things.
I'd prefer a K8s Operator or plugin (like Istio) to get all the pack.
Google operates at such a scale with a wide surface area of serious production considerations that even “ergonomic” solutions internally feel extremely heavyweight externally.
Source: I’m an Xoogler
I've been still just like, making VM's with proxmox, then putting my agent in the machine and letting it run free (with my dotfiles setup script making dev env pretty much free, though I could also just make a VM snapshot). What's wrong with that? Is that not the scalable solution for enterprise rn?
That said, I think Google's ADK ecosystem and this new AX platform is promising--I would expect Google to maintain this and other tooling around this for years to come.
To the Googlers out there: is Google using this at any capacity for internal projects?
The same Google that pulls plugs on a whim?
btw its the same google that has already killed its "gemini cli" and re-introduced it in the form of "antigravity cli"
Gmail for Your Domain/Google Apps for Your Domain/Google Apps/Google Apps Premier Edition/Google Apps for Business/Google Apps for Work/G Suite/Google Workspace
Do you see what you wrote?
guilty as charged
First time?
"Gosh, that Italian family at the next table sure is quiet"
IMO, you want the flexibility to create either: (a) permanent devbox VMs, and (b) per-task VMs
Agent sandbox platforms tend to be tuned for the latter, which sometimes involves VMM hackery for fast boot, snapshotting VM filesystem and RAM, etc.
Some workflows are a lot simpler if the multiple agents share a VM. These are workflows where agents must share state. A simple one we have: making related changes in our public OSS repo and our private repo, and then testing the change.
And other times you want to split up the tasks onto isolated VMs so they don't interfere with each other (ie run two dev servers without database or port collisions).
I tweeted a bit about this (https://x.com/dbmikus/status/2099264325231771878) and had a little debate with folks about ephemeral vs persistent VMs for agents
That said, most of the time, I only want an agent to work in one repo. I could give it multiple repos at once and instruct it to work in just one, but that risks it forgetting my instructions and it can load more context into the agent's window.
I think my main point was to have flexibility about the topology of VMs, repos, and agents.
Now the stuff people are coming up with is: how do you do authorization in this model? do you need a full sandbox all the time or can it be a workflow? how do you specify an agent is it a prompt or does it have some kind of control flow structure? How do you coordinate among many running agents?
I would say thats where we are now is there’s loads of people all solving the same problems a bit like when CoreOS, Kube etc. were all competing.
Code execution (usually TS/JS or Python) is useful most when you are dealing with truly open ended problems. It's the opposite of the use cases of most enterprise SaaS.
What's the actual realistic threat model for median developer or median user here?
By realistic, I mean that leaking your grandma's recipes or your SSN or your million dollar idea to some pastebin is neither likely nor going to meaningfully make things worse for you, or be useful for any malicious actor. Surely this is not what everyone is worried about?
I used two agents: One with network access to collect the evidence, and one with everything except the model endpoint cut off, which did the analysis.
The second agent's entire input was attacker-authored. So PHP droppers, obfuscated loaders, database rows, filenames, blah blah.
In this case I'm more worried about hostile input attacking the agent, and I need to contain the damage. My sandboxing solution does that by restricting access to the source data, making it read-only. The work dir can only transfer data via patch and apply (like a git workflow), so even my workspace can't be modified until I approve each change. And then restricted network means that any compromise ain't going noplace.
The second agent couldn't even install PHP or contact any CVE site to check if it was looking at a known attack, and that was by design. All it could do is write up a report about what it observed, not make assumptions about what it is. I could then take its (much smaller) clean output and pass that to a third agent with network access.
This is forensic work, so of course not your median dev's bread & butter. But the attack surface is only just starting to be plumbed. Compromising input can turn your agent into their agent, planting things as easily as planting worms was back in the early internet days when people connected without a firewall.
You only really need to sandbox when you provide access to tools that are almost impossible to filter correctly, such as a bash tool or a tool for arbitrary code execution.
I've been working on https://lullabot.github.io/sandbar/latest/ which works with Proxmox for VMs (and lima for locals or regular linux hosts over ssh). There's a diagram in https://lullabot.github.io/sandbar/latest/why/#recommended-w... with what we're currently recommending. Though, after some feedback, I'm in the process of integrating a colleague's web-based review tool as it turns out many preferred fully reviewing locally instead of using draft PRs.
It's got some opinions in terms of default tools for our team and industry so it may not fit yours. Forgive some of the AI-isms in the docs, I want to get the UX and feature set to a solid place before doing a full review.
A little git log, show, diff almost always in another terminal.
A customer of mine wants to standardize his developers on a Jetbrains IDE for python but I think that he is late by one year. Furthermore he is using the subsidized plans for Claude, not paying for token, so it makes sense to keep using the Claude TUI.
Not sure what all the other folks are doing, but the industry/ecosystem tends to over-engineer every single thing instead of just working on the thing, while I just want proper code, proper design/architecture, and proper high-quality results.
Yes, but, wrong layer here. Giving the agent a computer use (a la bash) is what folks are after. A temporary sandbox with lots of control knobs and security bits is how you do that in (as you noted) an enterprise.
The problem is that you'll end up wanting to run 2 or 3 (or 20, 100, 10,000) agents at once and that gets very hard with a single VM.
There's also an argument that you should be using a separate sandbox for each code operation a LLM performs (or at least each set of related operations). That's even harder to do with conventional VMs.
VMs are better for personal assistant work, GUI clicktesting, investigating bugs in your personal dogfooding dev instance and anything you haven‘t yet made repeatable and fast to set up.
Sandboxes are better when you need resource isolation or security and have a graph of tasks to work through. My agents often starve each other on one VM, so if they don‘t need any of the above it‘s just easier to isolate them.
Everyone is working in this area, including me [0], but either option really isn‘t that convenient to use yet. It‘s a bit of a „isn‘t Dropbox just FTP on a VM“ moment right now.
Since this is not the first mention of Dropbox I've seen in HN threads in the last 48 hours:
Let's not forget that Dropbox was at its best when it was "just" a streamlined ftpd over sshfs or whatever - when it was just "a folder that syncs". That didn't last long, the downfall started with them killing their most useful accidental feature[0], which started them on a path of enshittification[1], which they followed swiftly and diligently into complete irrelevancy they enjoy today.
So if the agentic tooling is now enjoying its "Dropbox moment", I implore people working on these tools, don't overdo it.
--
[0] - The "Public" folder initially supported direct linking, meaning you could publish static web sites by simply putting them in Dropbox/Public/, you could update the files there and changes were immediately "live". Notably, this was the heyday of phpBB and similar discussion boards, back between the rise and subsequent fall of free image hosting - so the ability to put images in your Dropbox/Public/ and hotlink them in a discussion was extremely useful and popular way to use the service.
[1] - They didn't just kill direct links, they replaced them with what I consider to be OG enshittification pattern - captive page that asks you to press a button to download. Yes, same one every "synced drive" service offers now, to enable various functionality that's 99% harmful to the user with the link.
Either way, this was the peak of Dropbox; after shutting down direct Public/ links, it was still useful for its main job as seamless cross-machine, cross-platform "folder that syncs", but gradually lost market share as OneDrive and Google Drive became more broadly useful (and had the advantage of being first-party on their respective platforms), and then Dropbox the company itself lost focus and tried a bunch of failed pivots in the direction towards cloudification, away from "just syncing files".
End result for end users? We now have zero options for bullshit-free, file-first, seamless "folder that syncs" experience for non-tech users (techies that like fiddling with things have Syncthing). Only cloud-first options remain, and they're full of footguns and enshittified to the core (which becomes apparent the moment you want to share a file outside of the vendor's cloud ecosystem).
Nothing at all. You'll know when you've outgrown it.
> what the general workflow is now that people are converging to?
Graph-based workflows where agents pick up work as it becomes available, structured output, while you manage the work queue and outcomes. Maybe? IDK really, it's all moving quite fast.
Where are the revolutionary software products?
Revolutionary products depends on revolutionary ideas, not faster execution.
I think that if they're unable to sell their idea well enough to find a co-founder odds are the idea would die on the vine whether or not they're able to implement it.
Works pretty well for me but I haven't put any effort into promoting it
- Do multiple tasks in the same context window / session, conflate different changes into the same prompt
- Repo mixed with old markdown files from previous tasks, excel and word docs and 300 playwright screenshots
- 5 tools all calling each other, test and deployment scripts are all markdown skills
Personally I prefer a ticketing system and isolated work trees
Probably we'll converge on a virtualised IO / Storage layer running microVMs beneath for isolation and security. Keep the network and storage layer separate for compatibility running a variety of stuff and a second security boundary.
As a test, I built a sandbox with only the host-side filtering proxy allowed for networking. 99% of traffic was HTTP. No QUIC at all.
npm, pip, apt, go, curl and git-over-HTTPS all worked on the standard proxy environment variables alone. No mirrors or other coaxing needed.
DNS is disallowed through the chokepoint, but that's no problem because the proxy resolves host-side anyway.
I outgrew this when I wanted to bring different sets of skills and templates to different machines, wanted to be able to share a small number of credentials, different agents in different machines, different egress rules etc. I wrote https://github.com/pjlsergeant/byre which gives you a TUI and some machinery for doing this easily on top of Docker or Podman.
We started building that but it quickly turned out to be too narrow. Often we want agents to do task that have no input ticket and often the output is not a code change (Slack bot, incident investigatior, scheduled daily tasks, ...)
No, giving the agent access to every single command on the system is not minimalist. It is actively detrimental if you want to do more than just attended coding with the agent.
Especially with autonomous agents, it's the only way to sanity.
We might need new OS abstractions.
While I feel like I have a decent understanding of the model landscape I'm feeling a bit lost at which agentic harness to leverage for local models. Hermes, Cline, Aider, Qwen Code, Goose, Pi, OpenCode, something else? I live in the terminal so Desktop UX is a bonus but not a must have.
Can I modify the antigravity settings/program to point to a local model? Where should I spend my energy?
Only complaint is that connecting the agent harness to my local model took more work getting configured right than I'd like, but that's been true of most harnesses I've tried as well. Most assume you're using a cloud model and local model configuration is a bit of an afterthought.
In case of omp, not sure if it's already at the node module package but you can just grab it from the links I shared and set it up.
* commands run by me (! prefix) are also sandbox blocked
* agent has no way to _request_ unsandboxed execution (e.g. if `kubectl whatever` is rejected by the sandbox, the model should have the chance to request permission)
* does not understand shell composition patterns (e.g. if `git status` is allowed and `git log` is allowed, then `git status && git log` should be allowed automatically)
* sandbox only supported on mac or linux. not both
All of that can be fixed by yourself. That's certainly the spirit of pi. But if you want strong defaults and batteries included (like omp promises) then that's just annoying.
- instruct model to write a markdown file with a phased plan to implement whatever feature or change I want
- start a new context, instruct model to implement one phase of the file
- review changes manually, then start a new context and have it do the next phase
- repeat as needed
I've never seen omp touch a file outside of the directory I start it up in, and the few times where I've been unhappy with a change git has been there to revert.
This could easily be a case of survivor bias but I've not had an issue with letting it go yolo yet.
I'm not a fan of open-core apps.
Always Pi.
Yes, it was developed by Google employees, that does not imply it has the full backing of Google, or Deepmind, or GCP. Notably, the website doesn't seem to claim this either.
A random example E.g https://github.com/google/filament#disclaimer
This is not an officially supported Google product.
https://cloud.google.com/blog/products/ai-machine-learning/a...
"Effort in GCP" is a red flag. (See Gemini CLI, which was shut down in favor of Antigravity CLI.)
AX is a layer that is closer to job orchestration, Agent Substrate. It's NOT an agentic framework. We use Antigravity for a few generative features but are abstracting away some of these components so anyone can bring their own. AX is trying to solve some tedious things everyone has to deal with. Layering execution with the underlying stateful worker, wiring up the network correctly, providing sub "task" identity, provisioning the right environment, proving stateful branching, auto discovery.
We want to keep Agent Substrate free of generative features and still need a layer above for some integrations that require agentic concepts and identity.
Google compute services are under GCP and I work on Kubernetes. So the comparison with Gemini CLI is not relevant here.
Plant many flowers, keep the ones that bloom and stop watering the ones that don't.
And anyone who bothers to just do a side-by-side feature comparison can immediately see how many features antigravity is still missing compared to Gemini CLI even today.
> This is not an officially supported Google product. This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.
We have used a few of Google's (smaller) open source projects, and in the last 2-3 years most of them are getting fewer updates if any updates at all. Some became very bad tech debt and we had to spend a lot of time migrating them.
Of course, that is the nature of open source projects (written in the license terms), and there is nothing to complain. But it's important to point out these days Google's open source project are not any more trustworthy than a one man's project in terms of support and maintainability. Personally I would stay away from them as far as possible. Especially if you look at what happened to Android, Gemini CLI etc.
(To be honest, even if it were officially supported by Google, that barely means anything. https://killedbygoogle.com/)
Google has around 200,000 employees. They probably haven't heard of most things Google releases.
The web page presents ax as a typical developer tool, but it's actually not for developers.
>A Model is not a model. It is a named model configuration: ...
Remember kids, a model is not a model.
If an agent is treated like nothing but a call to an external service (...which it is), everything fits in the existing programming paradigms. But I guess that's not very exciting. Only pragmatic.
I'm planning to buy a whole linux mini-PC to run my agents/code servers for more isolation. Codex/Claude Code let you run prompts on code over ssh (same with most IDEs) even on the desktop apps.
I wonder if that's going to be the new standard practice. You get a work laptop and an isolated agent box.
Running access control and network whitelists is always a maintenance challenge and it's easy to make mistakes.
I think you could get by with 1 computer, but it’ll have to have a pretty decent machine.
Between agents running tests, CI, docker image builds, an average $400 mini PC won’t cut it.
Don’t forget also many older mini PCs don’t support KVM. Some newer ones don’t support AVX/ mongodb.
It’s not so easy to buy any old hardware sadly.
A standalone machine is nice if you need more compute resources or if you want an always-on machine you can connect to from your laptop, phone, etc.
It doesn't look like Google's AX is quite the plug-and-play fit for running agents on a computer you own, since it requires setting up a K8S cluster, etc.
I think what's needed is something like a zero-setup combo of Tailscale and Firecracker
I'm trying to work towards that with my startup (https://github.com/gofixpoint/amika) but the bring-your-own-computer part doesn't work quite yet.
The project seems like an open-source initiative born out of the experience of some Googlers but not being used at Google. So, the title appears a bit misleading - people will be misled.