Posted by darkwater 5 days ago
GraphQL is absolutely terrible for human developers to interact with, but it's like Facebook could see into the future back in 2012. I cannot imagine a more perfect API surface for agents. With the REST API on GitHub, you can consume maybe 10 issue JSON blobs before your context window is blown out. With GraphQL constraining the results you can easily read hundreds in the same token budget.
Additionally, the # of requests your agents need to make can be reduced in many cases since GraphQL can join across types whereas REST APIs cannot. You essentially get savings in two dimensions here. Quota and raw token volume per logical response.
You could argue the rate limit guards should better reflect that, but that’s just not the reality of the system. Likely speaks to a lot of stability issues GitHub has been facing lately.
I'm not familiar with graphql but what would make something "invoke git", is it a technical thing or hyperbole?
So asking for the title of a PR might just be an extra column selected on a DB query. But calculating mergability status of that PR might be something else entirely.
That's why GitHub assigns "points" to each requests and deducts based on the data shape you request. For simple requests it's 1-to-1, but can quickly balloon
That also doesn’t help with load. So besides what agents are doing themselves the AI companies also have a responsibility of being good citizens.
As for GitLab, having hosted it for medium size organisations (~200 devs) and seeing how monorepo's work (they don't, we had GitLab's team show us that one page view made 50K db queries on our setup), please consult with your local admin team before firing GraphQL at it.
This is a solved problem. They just dump it into a file and `jq` or `rg` to find the stuff they need.
Agents are smarter than you think. They've been hill-climbing for generations in their RL environments.
The ones that get their context window blown out don't survive to launch
Replacing GraphQL and REST + OpenAPI OTOH I think is much more terrible. You have two API description language (one on the URL path, maybe the Zod schema, another one on the OpenAPI schema). Things like tRPC or magic functions are just using Typescript type system and comment to replace the schema that GraphQL already has.
The only thing I would bitch about GraphQL is it is quite hard to build an ad-hoc GraphQL server from the first principle, while REST is really KISS till the end. And GraphQL typically needs a lot more attention to N+1 problem.
We're building such an API for some of the biggest enterprises in the world. Many of them have very large (federated) GraphQL APIs across tens and hundreds of teams. From an agent perspective it's a lot easier to consume a single unified graph where a single query can span 5 relationships vs making hundreds of N+1 rest API calls across many heterogenous APIs from different teams that all look slightly different.
For our company, we advertise the graphql schema to bots and they can one-shot whatever task they're trying to do. I've found that it's so good that I cancelled building an an MCP server and any skill. Just a well documented GQL schema. It's pretty remarkable.
This gave me a chuckle, there's some subtle irony here- especially if the documentation of the GQL schema was work that needed to be done!
Most agents will use curl | jq to slice what they need (assuming a known API)
60/hour sucks. 5000/hour (a little more than one per second) is totally fine.
I'm chalking this up alongside Docker's decision to restrict unauthenticated pulls. Unauthenticated anything went the way of the dodo some time ago. If you want unauthenticated access, go run your own mirror.
If we had a system where people who access projects pay and popular FOSS developers get paid for it we'd have much better alignment.
My second thought was that bots would immediately try to circumvent such a plan. They'd probably spam Gitlab with fake repos to try to harvest those payouts.
People often aren’t hitting up GitHub directly to get or install open source projects: they’re going via homebrew, npm, or whatever, so the registry becomes the source. Of course you can install directly from GitHub even with package managers but most times you don’t and it’s increasingly seen as a security issue.
On your second point, ugh, yes, you’re absolutely right. I don’t think it would work exactly the way you describe but, if there’s some automated revenue sharing/distribution, you can bet that people will find ways to exploit it via some form of spamming.
Package managers tend to only need a small portion of what's in any given repo; some metadata to figure out what's going on and a single binary package is generally enough. The obvious answer is to separate those out and serve them differently from developers, who are actually monkeying around with the source code.
Normally, this is manageable, but for product categories with a huge network effect (like social media) this is a death knell unless you have the users of that product category already used to paying (Adobe's network effect driven business model comes to mind), and even then the switching friction is greater since that usually means many people paying for two systems for quite a long time, if not indefinitely.
Yeah I wonder if the math would shake out to make that make any sense. Each bot would require a paid subscription, so the only incentive for them to do this would be if there was some discoverability algorithm or SEO that that traffic helped push the content to real users
The guidance given seems to hurt open source projects, not help.
> Make the project private if the traffic is not coming from the audience you built it for, which stops anonymous callers reaching it at all. Or upgrade to Premium or Ultimate for much higher limits.
> What happens on October 19
> If you're close to a limit
> What changes and what doesn't
One request per minute.
https://gitlab.com/gitlab-org/gitlab
Makes 12 graphql API calls and 2 /api/v4 calls... So you get like 4 pages per hour unauth?
Yes, you get 64 more bits to make whatever addresses you want, but the prefix is still your fingerprint.
Browsing open issues or reviewing a few PRs will easily use more than one request per minute.
The limits are based on the average user but I wonder if the most common interaction is to view a readme and bounce.
I don’t know that putting a paywall up to learn from or even consider contributing to public projects is a good thing.
Rate limits, blocking, and pay-per-use are the only roads out and even those might not last as models get better at hacking and masquerading.
The internet we want to use LLM's with is simply not one that can support LLM's, and with LLM's not going anywhere, the whole experience of the internet is going to be forced into some radically less open and more expensive paradigm.
Policies like this just represent the beginning of the transition.
Loading https://gitlab.com/gitlab-org/gitlab is showing 14 API calls so presumably significantly more DB queries for a public page
It may look impossible right now. But what is impossible for real is to continue as we are. The damage that internet does to society is increasing by the day while its value is reduced (economic value, social value).
> you pay to access social media optimized to be interesting...
So is it commercial or non commercial?
Nothing is stopping you from creating a social network that is pay gated. Go build it. If you can't get anyone to sign up perhaps you'll realize it's not so easy as scapegoating addictive social media.
There is a whole cottage industry of people who legitimately make their living criticizing Facebook. It's a consumer software product. Yet few of these people seem to have their conviction extend to building an alternative that ever catches an audience. Why is that? Because addiction? Any other excuses?
Regulations, we need regulations for it to work. Capitalism is not going to solve a problem that capitalism has created.
> It's a consumer software product.
Please read: How Facebook contributed to genocide in Myanmar and why it will not be held accountable. - https://systemicjustice.org/article/facebook-and-genocide-ho...
> Why is that?
I value human live, I have morality and I try to be a good person and help society.
Who pays you to defend Facebook? Do you have shares? Do you have interests?
Is this the best you can do?
Except it's non-commercial, therefore valuable, therefore commercialized, therefore commercial.
You'd need a force strong enough to prevent it from falling prey to this tragedy of the commons, and that force would need to be stronger than the incentives to commercialize it. And that's where plenty of contemporary scraping-based salaries lay.
[0] One of many similar initiatives, I'm sure. Not an endorsement.
Your human engagement will attract said predators because it's a unique information signal.
Gating everything behind paid (but with no ads) likely would hurt a significant amount of lower income users.
The TOR network is probably the closest thing to what OP is describing but that is filled with illicit material and has no safeguards. We all pretend we don't like internet laws but those rules keep criminals below the surface and normal users from accidently stumbling on it.
Anonymous KYC is the way forward. Legislation on preventing bots presenting as legitimate users should be implemented. Similarly to how the world has mostly made robo calling illegal.
So, the precursor to online media has already gone through this paradigm shift.
Which is partially true, but it only shifts the distribution of the problem. Once your service gains enough popularity network effects cause it to gain value. You have to worry about high priced buyouts of the entire service (great for the site owner, terrible for the users).