Top
Best
New

Posted by Jach 16 hours ago

There's no reason for software to be slow anymore(danluu.com)
578 points | 415 comments
ehnto 14 hours ago|
One of the biggest causes of slowness is just waiting for web requests. The fact that so much software is either online or built using the same stack even if it isn't, puts all that software in this blocked/waiting state constantly while using it.

Anyone not in the US feels this even more since so much online is US hosted, 300ms for every little interaction adds up quick.

If your software has the affordance of a waiting dialogue or loading wheel for many of its UI controls, you are building with this default blocked assumption. Even if you are building something web based, ask yourself if that's actually necessary for your software or if you could build it differently to avoid constant UI blocking.

59nadir 8 hours ago||
This is an incomplete and quite superficial view of what is going on out there, in my opinion. I've worked on plenty of projects where the assumption was that since the round-trip to the server is going to take almost 100ms that'll dwarf anything that's going to happen on the server itself, justifying poor choices that lead to potentially adding a whopping 100ms onto that number. These numbers only get larger with a larger perceived "Nothing we can do about it" budget as well, programmers often feel justified in doing just about anything once round-trip time grows, not understanding that they're just adding to an already existing problem.

On top of that: Creating software that isn't outright wasteful in terms of performance isn't even hard, it's just a matter of not doing ridiculously dumb things. The problem I've observed in teams I've worked with is that the majority of programmers don't even know what the dumb things are, and wouldn't know how to even approach making something that's halfway fast.

Edit:

Unfortunately I think posts like these are only going to make the problem worse, because now people are going to ask for voodoo solutions to performance issues, when the answer to their problems was usually just "Maybe stop creating wasteful intermediate structures and just walk an array like a sane person" in 99% of cases. The first leg of any optimization journey in the average programmer's code will likely net tens or hundreds of times faster code, and that's actually all people were asking for.

The knowledge required to make those changes and understand them is fairly minimal, but the kinds of people who have to create spinners for webmail interfaces, have their application add 150ms on top of whatever round-trip you have for processing things counted in 5 digits, etc., have never bothered to even learn those things.

teeray 4 minutes ago|||
[delayed]
Yokohiii 5 hours ago|||
I don't think it's superficial, but two problems adding up. Previous poster is talking about general latency issues because everything is networked and potentially quite far away.

What you point out is slowness once you hit the entry point. Go, or similar languages, as a server language platform could have solved that problem from a computational perspective. But it did not for the most part. In my opinion people choose the faster stuff because it's cool and they have more wiggle room to cram in to get back to the slow status quo.

Everything is overengineered, software or distributed architectures, sound to naive human logic but alien to computers. It's an cultural problem, development is so deeply entrenched into "business logic" that the minimal viable and computational economic solution isn't even on the table. I don't even think it has to do with cost or feasibility, it's just that your random e-com manager wouldn't know what to do with you, if a programmer really starts talking about hardcode tech stuff.

jcelerier 12 hours ago|||
I have a new laptop with a rtx 5090. Opening any GL context takes more than half a second. There's tons of things that can be optimized and are pretty far from web.
fc417fc802 11 hours ago|||
I assume you're running proprietary drivers? Because I've never experienced anything like that on mesa. Launching an app that opens a window with a gl or vk context is so fast on my almost 10 year old hardware that it's nearly imperceptible.
adrian_b 7 hours ago||
There must be some quirk of whatever combination of software packages are installed there, but the proprietary drivers are not the culprit, at least not alone (i.e. there could be some interaction with other software packages with which I have little experience, like Gnome).

I have been using the proprietary NVIDIA drivers for more than 2 decades on various hardware, both desktops and laptops, mostly with Gentoo Linux.

Opening an OpenGL context or any other OpenGL operations have always been instant.

fc417fc802 6 hours ago||
So the usual self inflicted misconfiguration then.

In similar style I recently wiped a device that I thought had firmware that was slow to boot but it turns out that a hang and subsequent timeout due to something I had long ago misconfigured had been obscured by the previous setup that defaulted to hiding all details during boot.

userbinator 12 hours ago||||
We can only hope that the company building the hardware on which this new age of AI is based on soon starts "vibe-optimising" their own driver stack.
exe34 11 hours ago||
They don't want other companies to train on their IP and regurgitate it to random people.
genxy 12 hours ago||||
You should profile that, it is probably hitting the registry, the disk and maybe the network.

Try turning off wifi and see if it improves.

jcelerier 12 hours ago||
i'm on linux
4k0hz 11 hours ago||
No idea about your setup but that's probably a fixable driver issue. My desktop with an RTX 40-series GPU takes _maybe_ 4 frames to create an OpenGL context.
bcjones524 10 hours ago|||
[dead]
cavem0nkey 6 hours ago|||
Yes. Apple Music is the most egregious example of this. It could be ridiculously fast on your pocket supercomputer but the moment a web request gets fired off from stumbling blindly across the field-of-dung user interface, bam, you’re done. Especially if your network connection isn’t great at that time.

Things like this really pushed me to everything local systems. I’ll move actual files around if I want to do anything on the network. Or sometimes even use cables! Shock, horror!

sgarland 3 hours ago||
Sonos’ app is another example of this. As a very brief tl;dr if anyone isn’t aware of it, Sonos is a wireless speaker company that can group speakers in different rooms into zones, so you can have different music playing in different rooms, or all the same, or at different volumes, etc. The quality isn’t going to blow away audiophiles, but IMO they’re legitimately great.

The original design had the speakers setting up a private mesh network, and the app would send commands directly to the speakers via your LAN. Then, they got the brilliant idea to route commands via their cloud service. The app would send commands to an endpoint, which would send them back to your speakers. Imagine trying to smoothly fade volume with a WAN hop. This went over as well as you’d expect, and they’ve since promised to work on performance. Thus far they seem to have been doing so; it isn’t as snappy as the original, but it’s quite a bit better.

BorisMelnik 17 minutes ago|||
wow really? is this only for newer speakers or do older gen speakers work this way too
ryandrake 1 hour ago|||
So many developers do this, and it's infuriating. I have a device sitting there on my perfectly good LAN, yet if I want to remote control it, the brilliant software decides to send the commands to the Internet, then back to my device, then the response gets routed to the Internet, and back to my phone.

Device developers, stop doing this! You people realize that LANs exist, don't you?

cavem0nkey 5 minutes ago|||
As I said at [1] they want to centralise it as a control point so they can monetize it.

[1] https://news.ycombinator.com/item?id=49376040

fingerlocks 31 minutes ago|||
I was doing firmware + mobile a at large-ish startup a decade ago. Same story as Sonos; there was a push to go all cloud instead of our local network implementation that worked great.

I argued breathlessly against it for days. I’ll never forget the sales chad raising his voice to shut me down with a cop-out:“This is the way the industry is going!”.

It’s not the developers making these changes.

devin 10 hours ago|||
This is kind of a shallow assessment of "slowness". Slowness is a feeling, not a fact. Network is slow as a rule relative to other parts of the stack, but it is not usually what contributes to the feeling that your software is slow. It takes a good amount of incompetence and arrogance to cultivate that particular experience.
ehnto 5 hours ago||
Fair, it's shallow because I was succinct however, I do understand the problem space more in depth than this.

But if I were to pick one single thing that would speed up the most UIs across the board, it would be poor handling of the UI in networked systems. As you noted, that doesn't mean eliminating them, it means handling the inevitable in a way that doesn't tank the UI feel.

tosti 3 hours ago||
True, dev environments are fast. One dev implements a wrapper with roundtrips, another integrates it into a UI and no-one stops to think if it'll have terrible lag in practise. They don't notice so even if there's a ticket it'll starve and end up WONTFIX.

Somehow I don't think I'm the only one who presses a button and when nothing happens presses it repeatedly until something happens, or I kill the app, or even power off whatever piece of shit computer I'm using.

mawadev 8 hours ago|||
I agree and see this as a side effect of a subtler thing. As people shall be replacable, software is designed and subtly mimics the organization's communication patterns. Features are cut up into the tiniest pieces with clear separation from the start (at least its claimed), and over time whatever change or feature seems overly complicated, won't be done or won't be done in a sane manner because its uneconomical.

You are starting to get software with ticket driven development layered around glue code for existing libraries. I see no problem with libraries, it is just the architecture and vertical understanding that leaves a lot of performance on the table, because refactoring insanity takes resources and a lot of talking, understanding and convincing to be done.

Doing this across big teams starts to have downsides. So one team doesn't have a particular use case implemented or understood it and does not want to support it and you run code to compensate for this.

Best example in a monolith case is oracle...

leonidasrup 7 hours ago|||
Oracle:

    "Do not fall into the trap of anthropomorphizing Larry Ellison. You need to think of Larry Ellison the way you think of a lawnmower. You don’t anthropomorphize your lawnmower, the lawnmower just mows the lawn - you stick your hand in there and it’ll chop it off, the end. You don’t think "oh, the lawnmower hates me" – lawnmower doesn’t give a shit about you, lawnmower can’t hate you. Don’t anthropomorphize the lawnmower. Don’t fall into that trap about Oracle."
— Bryan Cantrill
jdndnrtigk 3 hours ago||
[flagged]
Lucasoato 7 hours ago|||
> ticket driven development

I feel this is so much real. Like if there were two kinds of companies: the ones who deliver, they care about their product but most of all they care about their customers, the managers get their hands dirty and everyone pushes towards the same direction; then there are companies in which you open a ticket and wait for two weeks for something that should take 5 minutes, customers and product don’t matter because you’re focused in cost attribution and no body does anything if it doesn’t come in your JIRA board, the managers are all coming from consultancy companies and all they do is finding someone to blame.

mawadev 7 hours ago||
I could go on about this all day. Best meme is when you estimate stories for storypoints and then they haggle with you without changing the content of the story. Meanwhile points mean actual time. Then on another side if you haggled down points, you then have more slots for more points. So you end up getting assigned 3x work for the same time. The blame game starts when the sprints elapse lmao
ck45 2 hours ago|||
Storypoints do not mean "time". Time would be too concrete and measurable and would be bad for selling more agile coachings. Instead storypoints represent "effort". What does it mean and how can it be used to estimate a shipping date? I was told that I just don't understand.
Lucasoato 5 hours ago|||
The only ways out of these situations are: a manager that understands what you’re doing; knowing that these “story points” are just indicative and there’s no broken incentive towards gaming them.
inigyou 4 hours ago|||
To dogfood this, rent a VPS in Australia and put a test environment there. Should be about 200ms ping or a bit higher. How it runs for you is how Australian users are seeing your site, even if their last mile connection is fiber.
bob1029 8 hours ago|||
The UI threading model is usually not the problem.

The problem usually comes from inappropriately arranging the systems of record such that information needs to be communicated beyond the scope of one computer in order to satisfy a single logical request.

Moving information between physical processors tends to be significantly more expensive than local computation over that same information. JSON serialization is a really good example of this. You need a network with bandwidth in excess of 10 GbE to begin overtaking simdjson.

SSR or SPA doesn't really matter if the server still takes a minimum of 300ms to compose any kind of response due to how its database or other infrastructure is set up. Information theory doesn't care how clever your loading indicator is. If the information isn't available, we can't do anything meaningful. Stringing the user along with psychological tricks is a lot cheaper than hiring a skilled developer to do it the right way.

inigyou 4 hours ago||
Networks are much faster than you think, it's networked software that tends to be slow. 10 GbE is now table stakes, you should fire any vendor who can't offer it. I certainly don't require your internet connection to be 10Gbps, or all your desktop machines, but your internal server network should be if you're building a new one in 2026, because there's no excuse not to any more.

> Information theory doesn't care how clever your loading indicator is. If the information isn't available, we can't do anything meaningful.

Yes we can - we can fix why the information isn't available. If someone said to you "sorry, we don't have the info because the other thread is doing Sleep(5000);" you'd call them an idiot right? You'd go and delete the sleep call to make it faster. Most real problems are harder than that, but there's no fundamental rule saying your database has to be slow. Ping time across your LAN is probably under a millisecond, so where are the other 299 milliseconds going? Is your database doing a full table scan? Is it using spinning rust for frequently accessed data?

sgarland 3 hours ago||
The majority of companies are building their stuff entirely in the cloud, where network speed scales with the instance size. I have had to explain this to multiple engineers at multiple companies, who are surprised to learn that network bandwidth isn’t unlimited.

As to your database comment, IME most of the time the bottleneck is the ORM and/or language. The amount of work an ORM does to generate a representation of a row is frankly shocking. Not understanding the cost of context-switching is the language half of it: Python, of course, is single-threaded, but you can use greenlets to cheat, because they’re I/O bound — except for all of them serializing behind a single process handling serdes for the queries.

alightsoul 11 hours ago|||
In my experience the main driver of latency is not ping time, but how long the server takes to process the request.
andersmurphy 11 hours ago|||
I think a lot of the time when people say the network is slow. They really mean their backend is slow.

With a fast backend ~1-5ms response times (not even that fast). Streaming compression over something like SSE to keep your response sub 1kb packet (roughly an ethernet MTU).

With a push based model, pushing data to a user is half their RTT latency. They will only experience their full RTT on actions they trigger.

Now the network to you is distance to the server (not your rail/nextjs backend taking 400ms). Things like 4G and 3G are fine. The real problem is when you have such bad signal you effectively have no down or up.

fhub 10 hours ago|||
1-5ms response time is clearly hard for most real world endpoints.
fallingbananna 9 hours ago||
I don't know how hard it is. But I can certainly say there is no business inscentive for it.

When it comes to improving performance by a few ms, or implementing a new feature, business people will always choose a new feature, unless the current performance is unbearably slow (we're talking regular 1.5s+ wait times for BE response).

And it's not even a modern problem, legacy software written 20 years ago has the same latency than most modern backends from my experience.

swiftcoder 10 hours ago||||
> With a fast backend ~1-5ms response times (not even that fast).

Even though benchmarks suggest this sort of performance should be trivial, most real-world servers I have interacted with do not reliably managed to process a request, make a roundtrip to the DB, and return a response in <5ms

inigyou 4 hours ago||
But why is that?
troupo 9 hours ago|||
Reddit's reaponses are rarely above 400ms. And yet their frontend routinely takes several seconds to render that response
thbb123 8 hours ago||||
More to the point it's that the server has to retrieve and massage data from several docker services to retrieve the full context needed to process the request
TylerE 7 hours ago||||
A lot of that is ultimately ping time, too though. Like making a naive number of round trips to a database not on the same machine.
NoMoreNicksLeft 9 hours ago|||
How long the server takes to process the 42 requests ahead of you in line, or possibly how long the 318 poorly architected microservices take.
vidarh 5 hours ago|||
It's particularly egregious to see the number of apps that will slow to a crawl even for things that works offline if the network is down or slow.
ttoinou 5 hours ago||
Being offline available/fallback and offline first are two different things
vidarh 5 hours ago||
Yes, but my point is that it's an awfully shitty fallback if you need to wait 30 seconds for something to time out first.

It's one thing not to e.g. spend the extra time to ensure everything is cached and mutations are queued up. It's another thing not to do the bare minimum to ensure what is already available and working locally is gated on the network being up.

Case in point: The other day I was checking our train tickets in an app, and the network was awful, and the train tickets which the app has local copies of took 30+ seconds to appear when the network went down. Everything I needed worked once the timeouts had been hit, it was just ridiculously slow waiting for timeouts for functionality I wasn't trying to use to be hit first.

ttoinou 1 hour ago||
Yeah, those apps were all coded in perfect network conditions and the designers refuse to change the UX to inform user about origin of data (offline, last cached X mins ago etc.)
hsn915 14 hours ago|||
It's not very hard to engineer software with these two constraints at the same time:

* Must feel very responsive * Network requests can take up to 500ms end to end

otterley 10 hours ago|||
It isn’t, but at the same time, smart hackers were working with highly constrained PC hardware in the 1980s and early 1990s and were cranking surprisingly good performance out of it. Folklore.org has plenty of stories about it, and John Carmack’s early career history is very impressive. We mustn’t forget the demo scene hackers either.
inigyou 4 hours ago||
Recently I tried an approach of "just write the f*** code" instead of using infinite abstractions on a new UI side project. So for instance when you scroll, it shifts the pixels and just redraws the new exposed area. This is how stuff worked in the 90s. And it's blazing fast and uses very little memory. It's easy to mess up redraw code like that - in my case, when the window goes past the screen border and the pixels to copy aren't there. That's a bug we also had a lot of in the 90s.
alightsoul 11 hours ago|||
That's why phones and windows use animations. You can also use intersitials related to the product you sell. Users are usually fine seeing many changes on the screen quickly because it gives the impression that stuff is happening on the background. For example in the interstitial, use an animation that takes up a small portion of the screen and not just a simple spinner or loading icon. Something more complicated with 2 or more things moving or changing at once.
jbstack 10 hours ago|||
If your app absolutely must rely on the cloud for every one of its interactions, then fine. If not, you're just applying band-aids to a problem of your own making. Many apps could easily be local only, or local first. If you're not constantly accessing the network for information which could be stored locally, then you don't need to hide your app's slowness behind animations.
alightsoul 10 hours ago||
Canva replaced PowerPoint. Canva is cloud based and PowerPoint is not. There's so many apps that are cloud only so that corporate it no longer has to manage installations and users don't need to ask IT for permission anymore. I think SaaS doesn't really work without the cloud, you could technically do what adobe does but why bother with app distribution and windows' quirks. Cloud based web apps are write once, run anywhere come true with no installation required. Cloud based is more convenient for the user and the developer, at the cost of app runtime speed
xstas1 10 hours ago||
I like this example. A lot of people would prefer to use PowerPoint because they can still use their files and templates they made last year even if Microsoft doubles the price of Powerpoint, removes features, discontinues the product, or goes bankrupt.
alightsoul 14 minutes ago|||
I have seen many companies migrate to gsuite away from Microsoft. It's perfectly compatible with excel, word and PowerPoint files unless you use macros
inigyou 4 hours ago||||
Just nerds I think. I'm not seeing any evidence that people who want freedom from the cloud make up any sizeable market segment. Most corporations actually prefer the opposite - they prefer a monthly fee and continuous silent updates.
lazide 5 hours ago|||
Difficulty - Microsoft retroactively cancelling lifetime licenses.
inigyou 4 hours ago||
Let them sue you if they want to.
lazide 2 hours ago||
They literally unlicense the product out from under you as part of windows update. You’re the one having to sue, unless you never connect your machine to the internet anyway.
ddejohn 10 hours ago||||
> Users are usually fine seeing many changes on the screen quickly because it gives the impression that stuff is happening on the background

This gave me a chuckle because I personally hate things like watching the browser jump through 50+ redirects when logging into a website.

alightsoul 10 hours ago||
My bad, I mean the interstitial. Like it should stay for at least a second to not make it jarring. People believe computers need to think so you can't make things too fast either. Not the interstitial and not the app either, to the point you sometimes have to deliberately slow down the app, add latency to make people trust it because it "gives the computer time to think"
jtari3333 7 minutes ago|||
Literally the first thing I do is disable all animations when setting up an os
inigyou 4 hours ago||||
Sometimes. Not by default.
lazide 5 hours ago|||
Power users absolutely hate this mindset. It’s resulted in Apple animations taking seconds for something that was done before the animation even started.
duskdozer 10 hours ago||||
I don't doubt that was an original justification, but most of what I see are not for this purpose. Most of the time they're just adding unnecessary delay and CPU cycles.
lazide 5 hours ago|||
I absolutely hate this shit, because my goal is actually doing something in a reasonable amount of time.
zanderwohl 9 hours ago|||
This is why my most recent website does its server-side rendering on the client. I'm not even being that sarcastic. We stream a subset of the user's data (1 mb at most) in the background and have a WASM client that has the exact same server views. On SPA events, the WASM blob intercepts a lot of requests and can instantly render. Makes a laggy connection feel pretty quick.
roncesvalles 8 hours ago|||
>SSR but in the browser

We've come full circle.

K0IN 4 hours ago||
I think this is our third lap.
inigyou 4 hours ago||
Are you counting mainframes with dumb terminals then smart terminals then PCs, early SSR websites then web 2.0, the multiseat/thinclient craze of the 2000s then computers becoming cheap enough to stick them to the back of monitors?
philipallstar 8 hours ago|||
Why aren't those views just part of your SPA in the first place?
hiAndrewQuinn 2 hours ago|||
Unfortunately, you either put all the commercially interesting bits on your own server and let your customers eat the latency, or you ship it to the edge and let piracy decimate your profits. I don't think there is any technical way out of this, and probably not any reasonable legal ways.
Tanoc 1 hour ago||
I feel like you're hitting at the real reason why so much of this happens. Not necessarily the piracy, but centralized control of the data and data flow. Even if it's slow as hell to send off that data to proprietary company servers (or rented cloud servers) to verify it, you're still verifying it. You, as in you the company, can't do that if it's local compute only. You can't control whether the user installed your paid plugin or some free alternative. You can't control someone stripping out libraries or code to remove intentional friction points designed to annoy them into a higher tier of the software. You can't control whether or not they update, or whether or not you can force the software into end-of-life with an update despite it still functioning.

If there's a connection to your services outside of the user's machine you can control all of that.

ntoskrnl_exe 10 hours ago|||
Anyone not in the US feels this even more since so much online is US hosted, 300ms for every little interaction adds up quick.

Don’t know about that, pretty much everything is hosted on Cloudflare, Azure or S3 these days and all of them have at least one CDN on each continent.

taeric 10 hours ago|||
There are CDNs everywhere, true. It is not true that everyone deploys to them all, though. Also it dodges that the networks available to everyone are still not equal.
lancebeet 4 hours ago||||
I think that "almost everyone" has a multi region cdn, but fewer have multi region application deployment, or cdn workers handling a significant portion of the application logic. My experience here may be incorrect or not generalizable, but I've rarely seen web apps that are slow due to latency loading static resources, but I've often seen slowness from high latency of the API calls and due to large static payloads.
FridgeSeal 4 hours ago|||
That only helps the assets, does nothing when they run their actual backend servers in one of the US aws regions and your traffic has to traverse the planet anyway.
stymaar 7 hours ago|||
> If your software has the affordance of a waiting dialogue or loading wheel for many of its UI controls, you are building with this default blocked assumption. Even if you are building something web based, ask yourself if that's actually necessary for your software or if you could build it differently to avoid constant UI blocking.

I wish Atlassian listened to you.

bambax 9 hours ago|||
UI blocking can make sense in some situations; otherwise things "happen" suddenly that are unexpected.

I'm making a simple plugin for Gimp that sends the active layer to a model with a prompt; on Macs by default the UI waits for the request to come back or timeout; on PCs it doesn't, so the modified layer appears unexpectedly. The Mac experience is better IMHO and I will replicate it in the PC version rather than the other way around.

Jtarii 1 hour ago|||
Is the 300ms ping time why discord takes ~10-30 seconds to load?
bigtechennui 3 hours ago|||
There’s no reason for (most) software to just be hosted in the US anymore.

We as an industry should use AI to enable a standard of software quality that was previously uneconomical.

kilroy123 3 hours ago|||
One of my all-time favorite software quotes is, "The fastest request is no request at all."

After traveling around in places with very poor wifi/phone data speeds. I couldn't agree more with you.

giantrobot 1 hour ago||
This has been my bugbear for years. Even in places with good cell service on average there's a hundred individual places that have terrible service. Also a good signal to the handset doesn't necessarily mean good actual service. It doesn't even require traveling, just normal daily movements to get wildly variable network performance.

It's infuriating when it's obvious that the developer of an app only ever tested it in a simulator on their dev machine on their super fast WiFi. It never seems to connect with those people developing a mobile app (or web app) that the "mobile" part has a meaning more than just on a handheld device.

energy123 13 hours ago|||
Can't they use ML to predict where I'm going to click, and pre-cache the predicted page whenever the predicted button doesn't mutate important state? Or skip the difficult ML and have some basic rule of thumb that pre-caches frequent button clicks, using a markov chain, and conditioned on those pages being low bandwidth to pre-load.
markerz 12 hours ago|||
So Next.JS actually pre-fetches links when they move into the viewport or you hover over it. It's interesting, but then you get wasted battery on mobile while on bad networks. The world is full of tradeoffs. Tech workers tend to want to consume more battery and data to be faster. Other people want to do less work.

https://nextjs.org/docs/app/guides/prefetching#hover-trigger...

In my view, websites should not take seconds to load with gigabit fiber. Whatever happened to "mobile first"?

zanderwohl 9 hours ago||||
McMaster Carr website prefetches _all_ links upon hover. Saves you like half a second in many cases.
inigyou 4 hours ago||
segor.de goes one better and just downloads the entire catalogue when you first open it. About 2MB decompressed. Clicking and even searching is instant because it's all fully client side until you place an order.

Funny thing is it apparently predates JSON. It's a bunch of data[foo][bar] = baz; - go look.

(Website's in German obviously, and a surprising number of German electronic terms are very different from English. They use two different words for stranded and non-stranded wire.)

supriyo-biswas 13 hours ago||||
> pre-caches frequent button clicks, using a markov chain

I don't think the current crop of fullstack engineers would be hard pressed to know what a "markov chain" is, but in theory yes, you could emit a bunch of speculation rules[1] based on your predictions.

I should also say that markov chain based approaches have been used for fraud detection, e.g. identifying checkout anomalies by detecting the sequence of web pages that they clicked on, amongst other factors.

[1] https://developer.mozilla.org/en-US/docs/Web/API/Speculation...

spaqin 12 hours ago||||
The problem isn't in preloading, it's in how much data needs to be sent while quite probably most of the data could be either fetched on startup in an efficient format and rendered natively, or is completely unnecessary in the first place (telemetry, ads).
__MatrixMan__ 13 hours ago||||
If we stored the edges (links) and nodes (pages) separately, rather than requiring you to blindly run a node's code just to discover what its edges might be, then you could skip the prediction and instead pre-cache the next hop for all edges just in case you follow one. You could even do this to two or three hops.

This might seem wasteful, but if the web were content addressed instead of server addressed you could then be serving that cache to your municipality even after it became disconnected from the rest of the internet. Which sort of recasts it not like wastefulness but instead like fault tolerance and preparedness.

We could maybe even dispense with the servers entirely.

There are so many different ways to build a web. Why does it feel like we've landed on the worst possible one?

inigyou 4 hours ago|||
We have that, it's called an <a> tag.
__MatrixMan__ 3 hours ago||
Its not stored separately, so:

1. You need write access to the server if you want to add one

2. The server could change its behavior at any time and there's no way to know that caches now need to be invalidated

3. If something goes wrong with connectivity or name resolution, there's no fallback since the authoritative thing was not something durable like a trusted human via a public key but rather an ephemeral thing: a named server which has pinkey promised to stay online.

It asks the user to treat a server like a trustworthy source of perisisant data.

But there's no reason to couple these kinds of trust. The skills necessary to persist and traffick data are orthogonal to being trustworthy about content. Coupling them creates needless load on single sources of failure which are simultaneously single points for corruption to target.

Trust people, not servers. Use digital signatures to validate that what you're seeing came from those people.

<a> tags are the opposite of this. They encourage us to trust servers by name, which isn't really working out.

charcircuit 11 hours ago||||
>but if the web were content addressed instead of server addressed you could then be serving that cache to your municipality even after it became disconnected from the rest of the internet.

This is already possible without content addressing with CDNs. They can serve content from a local cache even when the host is disconnected from the internet.

TeMPOraL 10 hours ago|||
And the first thing webdevs did once this became widely available, is change their apps to cache-bust their code; between that, and the short release periods in webshit ecosystem in general, and security and privacy considerations messing up things as usual, the promise of users mostly hitting just local cache with any marginal request, never materialized.
__MatrixMan__ 2 hours ago|||
Without content addressing how do I know that whoever holds the cache hasn't tampered with the content?
charcircuit 57 minutes ago||
The integrity attribute of the <link> element lets you provide a hash to ensure the content has not been tampered with.
Onavo 12 hours ago|||
> If we stored the edges (links) and nodes (pages) separately, rather than requiring you to blindly run a node's code just to discover what its edges might be, then you could skip the prediction and instead pre-cache the next hop for all edges just in case you follow one. You could even do this to two or three hops.

Welcome to Next.js

__MatrixMan__ 12 hours ago||
Can next.js give me a page's links without requiring that I execute any code that I didn't have prior to visiting that page?

The .js part makes me think not.

pests 11 hours ago||
Speculative Rules API

  <script type="speculationrules">
  {
    "prefetch": [
      {
        "source": "list",
        "urls": ["/checkout.html", "/thank-you.html"]
      }
    ] 
  }
  </script>
__MatrixMan__ 2 hours ago||
Right, but my browser doesn't interpret that. The site tells me to run some code which interprets it.

I should be able to get the lay of the land without trusting the site enough to blindly execute whatever code it points me at. It's needless attack surface.

Also it's not really pointing me at data, its pointing me a certain kinds of requests which I have to trust will be responded to consistently. I'd much rather have a hash so if I have that data lying around I can just forgo the request entirely and use what's present locally.

victorbjorklund 10 hours ago||||
This exist but the downside is that it uses much more of your bandwidth and client resources (probably not matter in many cases but it does if on a phone in a country with bad connection) and your server resources (if not mostly static content)
DeepSeaTortoise 9 hours ago||||
Oh god, don't give them ideas. All ML is in-cloud AI now. I dread the day everything around my mouse movements needs to get tokenized and vibed into the ClosedAI cloud before my buttons start working again.
rypskar 10 hours ago||||
I would assume a big part of software have plenty of lower hanging fruits for speedup and don't even profile to find where the bottlenecks are
CalRobert 12 hours ago|||
Heh, this is how browser accelerators from the dial up era worked
graemep 7 hours ago|||
This is true, but it is an entirely different problem in a different place to what the article is talking about.

The optimisations the article is talking about would help even with this problem if backends responded faster - although not as much as actually avoiding unnecessary network requests in the first place, of course.

giovannibonetti 5 hours ago|||
Shotout to PowerSync for enabling companies to go in the opposite direction and build offline-first apps. My company is a (production) customer, and we recommend it.
tujux 3 hours ago|||
Yup, this is the solution for a huge class of "slow" apps. More alternatives like PowerSync here: https://zero.rocicorp.dev/docs/when-to-use#alternatives
kobieps 4 hours ago|||
\m/ d-_-b \m/
jmtulloss 12 hours ago|||
I’ve been playing with this with software that needs to work with agents and also without internet at all (we’re serving construction projects that have limited access)

Obviously agent access goes away with internet failure but the state doesn’t need to… we use CRDTs and a virtual FS. There’s a toy-ish version of the harness at https://ourhearth.ai … if local first is interesting to you I’d love your feedback

saidnooneever 7 hours ago|||
there is different scales at which software is slow. this os one and definitely a pain in the ass. everything being online for no good reason other than to harvest user data. which it really turns out to be every time. (for good or bad purpose).

second is on a smaller timescale. so many features and crap that is never asked for and never used is crammed into software so the systems that execute it are just juggling pretty much dead code in and stale data in their caches all days long.

korijn 6 hours ago|||
Incredible that the top reply is a cop out.
benj111 4 hours ago|||
Surely the biggest cause of slowness is doing more stuff.

We don't turn faster hardware into faster programs, we turn it into more program. AI isn't going to change that. We'll just get even more program because the optimisation has freed up space for that.

Unfortunately most of the time, the more program isn't for our benefit. I note that by far the heaviest program I use is my web browser. The one thing I don't get to choose what code gets thrust upon me.

inigyou 3 hours ago|||
A browser isn't really a program any more but a platform for running other programs. Like how javaw.exe is really Minecraft, firefox.exe is really YouTube. Go to about:processes to find more detail.
benj111 3 hours ago||
The point is, if I want to edit a text file, I don't need to use eclipse. I can use something that hasn't added loads of features. I don't need to use whatever Adobe product, I can use some paint app.

If I want to watch streaming videos, I don't have a choice about how I do that.

Fine Firefox is basically a bloated YouTube app. That doesn't change the fact that it is inefficient (from the pov of my CPU) for doing that.

aaron695 13 hours ago||
[dead]
VCFundedGenYer 2 hours ago||
The peak of fast software was definitely Windows XP, Windows 7, and OS X Snow Leopard. I don't see us returning to that glorious era.

Recently I was frustrated by Windows 11's seeming inability to open a context menu with acceptable speed - right click an item in the taskbar and there is nearly a 1000ms delay before the menu appears. That is unacceptable.

When I need to run old software, I now try to the "minimum viable runner" OS - start with an XP VM and slowly move upwards if it doesn't work. Obviously I'll lock it down from internet access/etc., but it really shows that modern OSes really don't have a grip on performance.

mxuribe 1 hour ago||
As much as i dislike Windows, i agree that the versions of the OS from that era were really fast (except for maybe Vista)!! As far as OS X, I had used OS X back around ~2007 - 2010, but can't recall what versions it was...and it performed fine back then too. I've been on one or another linux distro since around 2003 or maybe 2004, and have been running linux as my primary driver for laptops and desktops since maybe 2010 or so(right after i uncoincidentally abandoned OS X). While any issues that linux runs into tend to mostly related to proprietary drives, the majority of the time, things run fast...just like Windows did of that older, golden period of performance - and many times much better! I don't say this to sway anyone to move over/start using linux...and, in fact, linux is still far from perfect! Rather, its to show that there is joy in computing that still exists somewhere in the world. I have found it in linux, but i'm sure others have found it elsewhere as well.
titzer 45 minutes ago||
I dunno, Ubuntu got bloated. I tried to put 24 on an old Chromebook at it was atrocious. Going back to MX Linux made the UI relatively snappy.
mxuribe 35 minutes ago||
You're not wrong! Some releases (and not just Ubuntu), it does feel like some distros are inching towards their "Windows Vista" moment, where UI aspects slow some things down...But, i think i tend to not get impacted by this since I often use KDE or XFCE....which is not to say that these are perfect nor the lightest-weights either....simply that, i know what you mean, but it doesn't tend to hit me much. Oh, and MX Linux i haven't touched in a long time, but yeah it used to fly and be super snappy back when i had used it!
5w6u56uw 1 hour ago||
The peak of fast software is right now within the open source ecosystem.
eaftan 13 hours ago||
I've been working on a similar agentically engineered regex project called SafeRE:

https://github.com/eaftan/safere

https://eaftan.github.io/safere-intro/

Mine is for Java and is intended to be production grade. The first goal is to guarantee linear-time behavior to prevent ReDoS attacks. My collaborator and I have recently been optimizing it to try to surpass native RE2 in performance.

It turns out optimizations are incredibly well suited for an agentic loop. You've got concrete acceptance criteria (must show a meaninging improvement on a benchmark case, must pass tests). The agent is really, really good at using tools like a profiler and disassembler, better than I am (and I've been doing this for 20 years). It also papers over things that would take me a while to learn, like how the in-incubation Vector (SIMD) API works in Java. I understand the concept but it would take me a while to understand Java's implementation. The agent can just read the docs and go.

The key is creating a good benchmark suite and ensuring the agent doesn't ship optimizations that are too narrow or too focused on the benchmark cases. You also need a really strong test suite to make sure you're not regressing correctness. SafeRE has billions of tests; a subset of several million run on CI, and the others run on-demand.

MattPalmer1086 5 hours ago||
Looks nice!

One little thing I spotted is you use Boyer Moore Horspool for fast literal search. This is actually not linear in the worst case, although it is almost always sublinear. Worst case would be a literal composed of the same character searching a text of the same character, where it becomes quadratic.

You can actually search strings with character classes using Horspool if you want to, and I have some enhancements to basic Horspool which could maybe help. My library, byteseek [1], implements these.

I also have a much faster algorithm, HashChain [2] which also has a guaranteed linear time version. This was published in the Symposium for Experimental Algorithmics in 2024.

[1] https://github.com/nishihatapalmer/byteseek

[2] https://github.com/nishihatapalmer/HashChain

MattPalmer1086 4 hours ago||
Reading a bit closer, you seem to track how much work is being done in Horspool on each character comparison, and then fall back to the linear KMP if the work budget is exhausted.

This will first massively slow down the Horspool scan, and then once you have done all that work, you rescan it all from the start with KMP if it is doing too much.

One little fix might be to only add to the work counter and compare it outside of the main character comparison loop.

But it would be better to use the linear version of Hashchain. It also uses KMP to make it linear, but it is fully integrated and you would not need to track the work or restart scanning at all. And its a lot faster than Horspool anyway!

za3faran 12 hours ago|||
This sounds very interesting. Which JVM profiler do you use?
stickfigure 12 hours ago|||
Not OP, but I went through this last week. I (ok, codex) optimized a hot path in some Java code from ~350ms to ~60ms, which made a substantial difference in "is this whole business going to work".

My Java profiling knowledge is... let's call it "antique". I was really not looking forward to ramping back up for this work. Turns out, I didn't have to do any of it. The LLM chose the tools (flight recorder) and even built a JMH (also new to me) harness to experiment with different algorithms.

About half of the optimizations were things that I would have figured out on my own; the other half were definitely "wow" moments.

The whole thing was done in a couple hours, with just a few back-and-forths. Sans AI, it would have taken a week, with nowhere near the same gain. I'm impressed.

"Figure out how to make this process fast" is really a perfect activity for LLMs. And the prompt doesn't really have to be much more sophisticated than that.

chii 10 hours ago||
> the other half were definitely "wow" moments.

do you have some samples? It would be interesting to learn what it might be.

stickfigure 2 hours ago||
Here are two:

* Using spherical points instead of trig to calculate distance between two geo locations.

* Packing data to minimize memory bandwidth consumption. Converting arrays of objects to multiple arrays of their component parts I sort of expected; bitshifting to pack and unpack multiple values into a `long` I did not.

Maybe other people would find these obvious, but I don't usually have to optimize at this level. My mental model of the relative speed of some CPU operations was a little out of date.

eaftan 11 hours ago|||
I've been using:

JMH as the framework to write microbenchmarks. It takes care of dealing with JIT warmup, etc. It's the standard way to write rigorous Java microbenchmarks.

async-profiler (https://github.com/async-profiler/async-profiler) for profiling. Java has a problem where many profilers are based on safepoints, which are biased toward particular program points. async-profiler is not biased in this way.

Java Flight Recorder for memory allocation data.

One thing I've observed in all of this is that it's really useful to have expertise in the programming language and ecosystem you're writing in, otherwise it's all Greek to you and you can't really guide the agent to do the right thing. I have opinions about e.g. profilers and I can point the agent to one that I think is more accurate than other options.

te_chris 9 hours ago||
This, 100%. I’ve rewritten some hydrology code in rust using codex sol to first a) profile and create comprehensive tests of the python, b) create full benchmark suites, c) create full scientific benchmark suites, then d) port to rust using a few different techniques.

It works, it’s at least 5x faster, sometimes much more, and memory use is like 10x less and even less in cases where lots of map tiles are involved.

This shit rules.

What I’m wondering now is can we reliably evolve python and have codex act as an extremely unreliable transpiler to the rust.

embedding-shape 7 hours ago||
> What I’m wondering now is can we reliably evolve python and have codex act as an extremely unreliable transpiler to the rust.

Why you even start with Python at this point? Just write the Rust version straight up instead of porting things?

Personally I used to use dynamic languages for most things, because development and maintenance is so much faster and easier, particularly for larger projects (granted you know how to work with those sort of languages), but now when the LLM writes most of the code, I'm able to work as fast with Rust as with I used to be able to do with Clojure or other dynamic languages.

te_chris 5 hours ago|||
Because it’s a port of a scientific process and that’s what the scientists work in.
embedding-shape 57 minutes ago||
Ah, sorry, I didn't understood you weren't the original author of the program you're porting :) Cheers for the additional context.
paganel 5 hours ago|||
What happens when LLMs will become either too expensive or unavailable?
scronkfinkle 5 hours ago|||
It's becoming increasingly evident this is not how things are going to shake out. Even without frontier models, running Qwen 3.8 27B has demonstrated for me and others a "good enough" competency at general programming. Additionally, large open weight models have proven themselves as viable alternatives and we're still in the early years of dedicated hardware
embedding-shape 57 minutes ago|||
I guess lots of Rust developers will find a lot of employment?

I mean what happens with network engineers when the biggest network of them all goes down?

mccoyb 14 hours ago||
Here's this boiled down:

> A stochastic search process with an executable optimization objective over space of programs S can only maintain or improve the objective

This is superoptimization. We've known this since the 80s (Massalin, STOKE is more recent: https://github.com/StanfordPL/stoke) The only novelty is that the proposer is now way better with LMs.

Further, there's a large number of reasons for software written by agents to be slow:

- LMs still don't do data or hardware-oriented design well out of the box, and therefore if you're engaging in any sort of serious novel work, beyond porting an extremely well-understood program with extremely well-understood workloads, you're going to be spending hours tracking down bad allocation decisions (c.f. why TigerBeetle doesn't use agents), which are often the root of evil (before you'd reach for anything further)

- The knobs you'd need to get serious performance are nearly unreachable in languages which LMs are good at (even Rust requires a discipline that the default language doesn't enforce). When you drop into the lower realms, you're trading consumption context for access to these levers. The levers are also "soft": you find yourself writing a bunch of skills, and tools to try and enforce the discipline.

The reality is to get performant code (quickly) out of an agent, you need to know how to write performant code (and you need to know how to surface the information that you'd use to create a verifier for such a thing to the agent), which 99% of developers do not know in 2026.

Sure, agents can teach you how to do this -- but it's one of these things where iykyk.

Experience: I've poured 10s of billions of tokens into Zig with the best agents and I have the time and space to try these things.

If you want to start learning the discipline, I'd recommend matklad's + TigerBeetle blog -- as well as hardware-oriented design.

Aurornis 14 hours ago||
> - The knobs you'd need to get serious performance are nearly unreachable in languages which LMs are good at (even Rust requires a discipline that the default language doesn't enforce). When you drop into the lower realms, you're trading consumption context for access to these levers. The levers are also "soft": you find yourself writing a bunch of skills, and tools to try and enforce the discipline.

Hasn’t been my experience at all. The latest LLMs can knock out assembly optimized subroutines and benchmark 100 different variations faster than I ever could dream of.

mccoyb 14 hours ago|||
That's fair for a well-scoped subroutine: what I meant is that if you ask an agent to write a compiler and let it rip for a few days, you are going to be spending a few more days correcting the default behaviors in the distribution, which often do not tend towards hardware-oriented design.

To correct those behaviors, you're going to write tools and skills, and that's going to help, but it is still clear that you are fighting the distribution (today).

embedding-shape 7 hours ago|||
Yes, say "Build a compiler" will require you to clean up stuff if you leave the agent for days, but not because of the LLM or the quality of the tool, but because you hardly specified anything, so of course it's gonna make assumptions you need to correct.

If you instead spend a day writing a proper specification, then ask the agent to spend a week implementing that, you'll need zero tools and skills afterwards to clean it up, because there won't be any misunderstandings, assumptions or other things, just code fulfilling what the specification says.

Granted, this does require you to not use obviously dumb models, like anything you can run locally today, and at least within reasonable range of SOTA models. But they been able to do this for 6 months or more at this point.

mccoyb 44 minutes ago|||
Bro, what the fuck do you think I’m doing? Do you think I don’t know about spec driven development?

What is the most complicated thing you’ve built with LM agents? Have you done it with a single spec? How novel was it?

This comment is so laughably “you’re holding it wrong” I can’t respond to you seriously.

> If you instead spend a day writing a proper specification, then ask the agent to spend a week implementing that, you'll need zero tools and skills afterwards to clean it up, because there won't be any misunderstandings, assumptions or other things

The set of software that has followed this process is measure zero.

n4r9 6 hours ago|||
You're making assumptions; OP made no mention of how detailed their spec was.
senderista 13 hours ago|||
Agreed, IME Fable can churn out decent SIMD kernels optimized for whatever tradeoffs you give it.
josephg 13 hours ago|||
> iykyk

A story.

I have a friend, who is - like me - interested in the CRDT / collaborative editing space. He asked ChatGPT to write him a CRDT. Then he grabbed every good CRDT implementation, and asked chatgpt to benchmark and optimise his CRDT, using tricks and techniques from existing hand-optimised CRDTs. He got massive performance gains by doing this - which is really interesting! I think it helped that he had a clear objective function, and chatgpt could look at other projects for ideas on how to optimise.

He proudly boasted that his resulting code outperformed my diamond-types library. I asked him if he was comparing against the native implementation, or the -Oz webassembly build, running in a wasm vm. It was the latter. When he tested it properly, his CRDT was - and is - significantly slower than diamond types. As far as I know, chatgpt still hasn't been able to catch up. I tried myself using fable. Even with reference to my source code, Fable still doesn't understand what I did in diamond types and why. (... Maybe I should document what I did!)

I think his technique itself is solid though. I tried it myself. I asked fable to write a custom binary serialization format & parser. Then optimise. Then optimise, with explicit reference to existing libraries. Optimising with reference to other code made a huge additional difference. It is now nearly as fast as those libraries. (But still not faster than them.)

My takeaway is this: I think LLMs are exceptionally good at reading and understanding code. If you guide them to do so, they're good at profiling and benchmarking. But it seems like they're not very good at coming up with novel optimisations. If you have an obviously slow program (for example, some slop claude wrote), you can often get big speedups by asking it to benchmark and optimise. But if you have a complex, already well optimised codebase, like the zig compiler, claude doesn't seem very good at figuring out novel ways to improve things on its own.

This is good news for the 95% of slow software out there. But bad news for the 5% of us who write fast code already, but want our code to go even faster.

jkercher 2 hours ago|||
I think you nailed it. And that would make sense and should be expected. 99.9% of the world's software (the training data) is several orders of magnitude away from max performance.

Another possible confusing thing for an LLM is that getting close to max doesn't necessarily require any "tricks." A big part of getting in the ballpark is just not doing anything you don't have to. If program A is faster than program B, most of the time is not some magic algorithm. It's that program A just did less stuff.

FacelessJim 7 hours ago||||
Small OT:

> Maybe I should document what I did!

Please do!

Me and a bunch of friends worked on a project that used diamond-types as the backing CRDT engine. It does indeed go brrrr. But damn, it took way too long to reconstruct what it was doing (we needed some more fine grained knobs, so we were playing with the frontier directly). We eventually moved out to something a bit better documented, which was a real pity. I really liked the general architecture and simplicity (of the text-only based version at least).

josephg 7 hours ago||
Thanks! If I can ask - where did you get stuck?
gritzko 9 hours ago||||
Hi Joseph,

A story for a story. I had my CRDT implementation in libdog, which does per-token CRDT weave/diff/merge over a DAG of git blobs. It was written by Claude 4.8 I believe, in several iterations. It was, as you may guess, a piece of neuro-slop that passed the tests by some miracle. Once I had some time to look into it, I used a trick: I supplied it with my article on Chronofolds and some helpful kicks in the butt. It implemented everything correctly on its k-th attempt, k<5. Then I used it with full intensity for three months without thinking twice. Now I have started mass-using it to resolve permalinks in the code. Like, tens of files to chronofold-ize per one commit. It is now showing up in the profile, so I may look into it once more.

Conclusion: it sort of expands its context and prompt by association. It has no sense of direction of its own maybe. Once you know what you are doing, you can ride it fine. If not, it makes a misstep so later things go haywire, and you are left guessing why. (And how can you guess if you have not had the experience. We all grew with Commodores and suchlike. I recall Spectrums, Robotrons, and Poisks. My friend sits on exam committees, says the youth arrives flatlined after 3 years of GPT. His words. Whatever.)

josephg 7 hours ago||
> It has no sense of direction of its own maybe. Once you know what you are doing, you can ride it fine. If not, it makes a misstep so later things go haywire, and you are left guessing why.

This is a perfect description. Last week I asked claude opus to get AAC audio working in davinci resolve on linux. It managed to add aac in mov and mp4 containers very quickly. But mkv was another matter. For mkv files, resolve doesn't use ffmpeg. Instead, it has its own parser. Claude got totally lost down a weird rabbit hole trying to add aac support to resolve's mkv code. It was really struggling. Claude even knew it was lost - it kept telling me we should cut our losses and I should just release aac support without mkv.

Eventually I gave it the executable for davinci resolve on mac, which has aac support. Claude found the corresponding part of the code for the mac version and used it as a reference. Turns out, claude had made some much earlier mistake. Just like you said, it was going down a wrong path. Then it couldn't stop itself, and it kept making it worse.

Using the mac version of the binary as a reference, claude figured out how to get everything working very quickly. But - I'm left wondering. Maybe my real mistake was using Opus and not Fable. I wonder if fable would have been smart enough to figure out the mistake and course-correct.

gritzko 7 hours ago||
Like Aladdin's jinni is a slave of the lamp, LLM is a slave of its context window :)
fc417fc802 11 hours ago||||
The refutation of your takeaway is autoresearch and similar. They can brute force novel optimizations (and generally achieve superhuman performance) when provided with an appropriate environment.

Of course that doesn't mean they have a human level mental model and associated novel ideas. Brute force can be effective but remains entirely unsatisfying from an academic perspective.

josephg 9 hours ago||
True! How do you set up auto research loops? Are there any special tricks to it?
fc417fc802 9 hours ago||
I don't know if it's the first and it certainly isn't state of the art at this point but I think karpathy/autoresearch is a quintessential starting point.

TBF as an earlier commenter noted it's "just" superoptimization using an LLM as the proposer so there's a lot of relevant prior art.

zem 10 hours ago|||
yep, i think the real LLM superpower is knowing that something has been done before and having access to the code that did it. so much of even novel software includes bits and pieces that have well-optimised existing solutions, and the bot knows those solutions a lot better than i do, and can even pattern match them from the general shape of the problem.
josephg 9 hours ago||
Yes. It’s also excellent at reading large codebases and putting together a picture of what’s going on. I’ve been using it a lot lately to brief me on projects and design decisions. “Look at these two projects. They both solve task X. Write a report about their similarities and differences, and the tradeoffs as a result.” And then I ask followup questions. Saves a ton of time.
hiddencost 12 hours ago||
> A stochastic search process with an executable optimization objective over space of programs S can only maintain or improve the objective

Reasons this doesn't follow: (1) Benchmarks never match real world use, and many optimizations the improve benchmarks degrade cases that aren't measured (think about how CPU cache behavior can be surprising) (2) In software performance optimization, frequently there is significant noise, from many sources. This makes it difficult to guarantee that a measured change is actually an improvement.

fc417fc802 11 hours ago||
Both of those things are indicators of deficiencies in the testing process.
joss82 2 hours ago||
That’s the exact opposite impression of my recent user experiences of software: slower than ever.

Contrary to what initiatives like the tigerbeetle team is doing with tigerstyle, or the 10 nasa coding rules, code created by llms tends to be verbose and slow.

p0w3n3d 2 hours ago||

  Lol got my mac M5 128GB I'll import a JS framework for multiplying numbers 
Careless coding has been introduced by people saying "programmer's pension is more than double the RAM" but it is no longer the case. The windows UI could occupy 30MB at most. But they chose differently
hunterpayne 11 hours ago||
"LLMs are causing slow, bloated, code are going to eat crow once they re-write everything in super-optimized assembly."

This person doesn't understand how to make efficient code. I can write code in almost any language (with a couple of exceptions) that outperforms "super-optimized assembly". Writing efficient code isn't about the language, and often isn't about the best algorithms either (but sometimes it is). Its about optimizing memory and cache use. And that's orthogonal to anything the author is writing about. Also, LLMs are terrible at optimizing memory utilization. There is just too little training code that does it well and far too much that doesn't.

As proof, I'm can literally feel the web getting slower and I bet many others feel this as well.

iainmerrick 3 hours ago||
I think you’re talking somewhat at cross purposes to the original article.

The points I take away are:

- Good optimization is difficult and slow work, hence expensive, but LLMs can do it so we should be able to afford it more often now.

- There’s always a risk of over-fitting to your specific problem, but if everyone is now making bespoke optimizations maybe that isn’t actually a problem.

You said:

LLMs are terrible at optimizing memory utilization. There is just too little training code that does it well and far too much that doesn't.

There’s probably something in that, but it can be mitigated by testing against a local benchmark. LLMs are good at iterating tirelessly and finding incremental improvements. And as noted above, it doesn’t necessarily matter if your benchmark isn’t fully general.

malisper 11 hours ago|||
> This person doesn't understand how to make efficient code

The author is one of the most knowledgeable people about performance there is

hunterpayne 9 hours ago|||
The author didn't write the quote.
Kiro 5 hours ago||
Just admit that you were wrong instead of this doubling down nonsense. LLMs are amazing at optimizing memory utilization. They have no problem obsessing over fitting as much data as possible onto a single cache line and micro benchmarking cache hits.
rrook 5 hours ago||
This seems like way too caustic of a reaction, OP is correct.

If you actually know what you're doing in $language, and you know how $language wants to emit the assembly or whatever, there's no huge advantage to just programming directly in assembly.

geraneum 10 hours ago|||
Seems like author’s main focus recently is AI and agents unsurprisingly, hence the suspicion. But it seems like he has a backgrounded in relevant fields in the past.
whatisthiseven 10 hours ago|||
My contradictory proof: I have been working on an old service with tons of performance issues, from server memory bloat, client graph rendering, excessive network requests, excessive repeat rendering, memory leaks, resource leaks, etc.

The app and service are measurably and subjectively faster. Because I chose to have the LLM focus on solving those problems. It obviously can. It described the issues in big-O.

It is a priority problem, as it always has been, not a knowledge or skill problem, like it always has been.

99954bb63ccc 3 hours ago|||
I think it's more about not doing unnecessary things. Like a like a saying I heard somewhere "a clever person solves a problem, a wise person avoids it". Not every good idea _needs_ to become a feature. And if you think it's that good, give users the option to turn it on/off and track that as a metric.
ddejohn 10 hours ago|||
The author was paraphrasing a tweet written by somebody else:

> The other day, I saw a viral tweet saying [...]

momocowcow 6 hours ago|||
I remember the co-founder of Anduril Industries being the author of this tweet!
userbinator 10 hours ago||
Also, LLMs are terrible at optimizing memory utilization.

I've already posted this elsewhere, but here it is again: https://news.ycombinator.com/item?id=49226923

A vibe-coded OS that runs on an 8088 with 256KB of RAM.

As proof, I'm can literally feel the web getting slower and I bet many others feel this as well.

"It's not the tool, it's how you use it..."

intrasight 14 hours ago||
I've been using computers for 4 decades. They have gotten no faster. The nuclear plant computer system we built in 1989 had to present selected screens in 1 second. I don't think any apps I use today can do that.
paulhebert 13 hours ago||
It’s all incentives. I’ve worked on web projects where the people in charge cared about performance. It’s easy to get sub second speeds if you start with that goal.

I’ve also worked on projects where the people in charge added 1mb client-side mapping libraries to render a static map and ignored my push back. Those website were slow

gwd 8 hours ago||
This. The author talks about the fact that now they can try far more experimental optimizations than they could before. But they already had an architecture with efficiency in mind, and were trying optimizations, to begin with. The kind of company that ships web UIs that take 5-10s to do some action 1) don't care or aren't capable of good architecture 2) don't care or aren't capable of doing the simplest low-hanging optimizations.
alightsoul 11 hours ago|||
At some point, probably in the 1980s, people probably decided that computers could update UI fast enough, so any additional compute power or speed has been used for other things like making it prettier or reducing development effort or time, and now ai
catdog 10 hours ago||
The problem is that the UI speed largely was not kept at that level but considerably slowed down again despite exponentially growing compute power.
duskdozer 10 hours ago||
I've heard it before and believe it that one reason is because many of the developers are working on new maxed-out machines and network connections both at work and home, so they don't notice problems for older or cheaper ones and/or can't justify it to management.
locallost 6 hours ago|||
What Andy giveth, Bill taketh away.
intrasight 6 hours ago||
In my case neither Andy nor Bill was involved. For that project we used Sun 3 workstations which were powered by Motorola 68020 I believe. So a more fair comparison would be that 1988 workstation against a modern Unix workstation. I bet the results would still be that a modern workstation would struggle to update the display in one second.
locallost 5 hours ago||
Well I don't know about specific examples, but it's a phenomenon that has been observed in many areas over decades, thus the funny saying.

After crossing 40 years of age, and working for a while now, I believe it's also because of politics. You might think the goal is to deliver the best possible product, but territory grabs within companies are important, and done by people that don't have enough skill other than territory grabs. E.g. look at Trump and his behavior. No skills other than having his way, and then he gets to decide.

Of the technical reason, for sure we underestimate how much faster technology gets. Fred Brooks had this example in his Mythical Man Month book, how the os/360 got the option of a disk drive instead of mag tape, but the result was worse because everyone assumed it was much faster than it really was.

cyteeditor 11 hours ago||
[flagged]
chvid 10 hours ago||
ChatGPT MacOSX is the only software that regularly crashes on my machine when its memory consumption for no apparent reason spins up towards 50 GB.

And that software is build by some of the highest paid software engineers on the planet with full access to all the LLM compute in the world.

blfr 10 hours ago||
Yes, but it's also built by people who would rather tinker with AI than build MacOS apps. Intrinsic motivation is very hard to beat, especially in subtler areas like good UX or performant software.
aureate 9 hours ago|||
OpenAI has 4,500 employees and a post-money valuation of $852 billion at its last funding round. You'd think they could get someone who specialises in writing MacOS apps to write their MacOS app.
jmalicki 9 hours ago|||
They can probably hire them, and they'd get PIP'd for not jumping when management says jump.

Toxic corporate culture can go a long way to self-sabotage.

chvid 9 hours ago|||
No. They want the AI to do it.
inigyou 3 hours ago|||
The genius play is to have an OSX expert write it with AI assistance, then let the press release say AI did it
oblio 7 hours ago||||
Hopefully at some point gravity pulls them back down and they reach the maturity and humility to realize that LLMs are just tools and even though they're good for many things, you can't do absolutely everything with them.

I think the "maturity and humility" phase of the LLM hype cycle is probably still 2-3 years into the future.

And that phase probably comes immediately after the "LLMs overinvestments have caused a massive global recession".

cpursley 8 hours ago|||
So have the AI build a hybrid Crux, Windows and web app with shared Rust base.
chvid 9 hours ago||||
I think the software out of the big AI labs: OpenAI's ChatGPT MacOSX app, Anthropic Claude Code - is a sign of the future to come.

Big, feature rich, built in quick iterations - but also, in particular if you look under the hood, of extremely poor quality if measured by traditional software engineering standards (code structure as exemplified by the leaked Claude Code source code, resource usage, "buggyness" etc.).

the_gipsy 4 hours ago|||
They absolutely do not have AI-heads with PhDs writing the apps and general infra. I've met a guy that worked at Anthropic, and he was "just an engineer", so exactly that: a top guy, but not specifically AI focused, doing regular stuff there but with infinite access to the top of the crop AI.

Now look at what software they make. It speaks for itself.

ponector 6 hours ago||
Coding is a solved problem. They simply don't care if the app is full of bugs. You are using it anyway, right?
jjcm 14 hours ago|
This speaks to me. I've been running an autoresearch loop the past couple of days to improve the load time of my various projects' frontends.

I've been really, really impressed with how effective this is. I went from a 4s load on simulated slow 4g to ~750ms: https://image.non.io/speedup-graphs.webp

Side by side vid of the results: https://video.non.io/speedups.mp4

This was for https://non.io, which is something I had purposefully written to be as fast as possible (hand wrote all the comopnents, didnt even use react).

I've been considering creating a skill / utility to do this based on learnings from the speedups - would others find this kind of thing useful?

ricardobeat 2 hours ago||
It still feels pretty slow, and I see some low-hanging fruit:

- in safari, every image is loaded twice, .heic and .webp

- default.png is re-downloaded 31 times, uncached

- images below the fold are request immediately, lazy loading would avoid that [1]

But most important, you have 229 requests for tiny files being served over HTTP 1.1. Without GZIP. From a pretty slow server - 700ms+ to download the main json data. Bundling your JS, or enabling HTTP2 or QUIC/HTTP3 alone would massively improve performance.

If this is the result of days of autoresearch, it's not really anything to celebrate.

[1] https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/...

Flamkuchlo 2 hours ago|||
It loads in 4 seconds for me on a fast landline.

And there are plenty of things you did not opimize for. Like making sure certain dimensions are already known to the dom renderer so that the layout doesn't jump around.

Or progressive images so that it doesn't just popup suddenly.

VladVladikoff 14 hours ago|||
> This was for https://non.io

Maybe hugged but feels really sluggish to me for it is.

throwitaway222 13 hours ago|||
Never visited before and the site was HN load speed. Very fast.
jjcm 13 hours ago|||
These are the stats I'm seeing: https://image.non.io/7a8adcb1-2a17-4e2c-b213-d0c3e173a1d5.we...

Should note the server is on USW and I don't have edge servers for it at the moment.

ivm 3 hours ago|||
It takes about 6 seconds to get the site fully rendered on MacBook Air 2017 in Chile. HN is slightly above 1 second in comparison.
roncesvalles 8 hours ago|||
>something I had purposefully written to be as fast as possible

and then you threw it all away by adding transition animation

ltbarcly3 13 hours ago|||
Why is it so slow? Like clicking around this is a very simple site, it seems like the fade in and fade out, besides being jarring and annoying, is just adding load time.
abanana 6 hours ago|||
An annoyance: clicking an image transitions the tiny image to a bigger size, then it gets replaced with a larger image file. The result is: the image slides bigger (but blurred), then immediately disappears, and reloads slowly from the top down. It's an annoying flash, and the slide to a bigger size was a waste of time.

That seriously needs some optimising. For example: on click, could the bigger image be inserted behind the small one so it's hidden; on load of the bigger image, hide the smaller image; then do the slide-bigger transition with both images together?

arvigeus 9 hours ago||
Autoresearch is crazy good! Closest thing to “Make this app fast!” we have now.
More comments...