Top
Best
New

Posted by razin 5 hours ago

RIP, vector database(turbopuffer.com)
238 points | 63 comments
gopalv 5 hours ago|
> This write amplification is large enough that our efforts to tune indexing throughput have started to hit diminishing returns.

> don't key on the ANN address. That is precisely the change turbopuffer v3 makes. As you can imagine, it is not a trivial change.

This is a direct parallel to how Postgres and Mysql built indexes.

Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table.

Postgres always points an index to a row-id within postgres which is an arbitrary value which changes on each update.

Mysql, always assuming the storage engine is pluggable, points to the primary index entry and adds an extra indirection to the lookup.

This means that you point the mysql index to a stable id, so unless you go update the primary key for a row, you won't have to update the indexes for all the attribute lookups you might have made to data.

I don't do databases any more that much, but the design for NIMBLE file format has a lot of quirks which are relevant to this specific idea (wide tables).

But the old Uber post about switching from Postgres to Mysql to prevent index amplification[1] is a direct mirror to this post.

[1] - https://www.uber.com/us/en/blog/postgres-to-mysql-migration/

malisper 3 hours ago||
> Your design choice went from a Postgres design pattern to a Mysql one. The difference is the reindexing cost vs the lookup cost - Postgres optimized for lookup and Mysql does for indexing on writes. Or more accurately, Postgres was better with good schema design using joins & mysql was optimized for a bad design with less normalization where many indexes exist for the same table

You are right that MySQL does better when you have lots of indexes, but I don't think the tradeoff is that the overall Postgres architecture is better with good schema design.

Having secondary indexes point the primary key enables things like undo logging, which obviates the need for vacuums - vacuums being the most painful part of Postgres. On top of that your primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear

tomnipotent 2 hours ago||
I think OP is just alluding to the fact that Postgres needs to do less work to go from secondary index to table data, since the tid is a direct pointer to the exact page and slotted entry while MySQL needs a b-tree walk.

> primary key index will be mostly cached so the cost of the indirection is much smaller than it may first appear

Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.

malisper 1 minute ago|||
> I think OP is just alluding to the fact that Postgres needs to do less work to go from secondary index to table data, since the tid is a direct pointer to the exact page and slotted entry while MySQL needs a b-tree walk.

Yes, this is true, but they framed this as "the Postgres approach is better when you have a good schema design", but that's not true. There are plenty of ways the MySQL approach is better even when you have a really good schema.

> Not sure I follow. If it's in-memory you save having to read from disk, but you still have to walk the b-tree to go from PK to data.

The point I was trying to make is that going to disk is going to be orders of magnitude slower than doing an in-memory B-tree traversal. Because of that, the cost of doing an extra b-tree traversal to find the page you're looking for is a relatively small cost compared to reading the page in the first place

barrkel 1 hour ago|||
MySQL was generally (pre 8) optimized for point queries on primary keys. So rows are stored in the PK index, the PK index is a clustered index. Everything more or less falls out of this.
phoghed 5 hours ago|||
> mysql was optimized for a bad design

TIL I should have been using mysql the whole time

woadwarrior01 5 hours ago||
Richard Gabriel's "Worse is better" vibes.
__s 40 minutes ago||
I literally had opening slide about Worse is Better when adding mysql support to peerdb. Lots of things in mysql have v1 as stupidest thing that works, then v2 fixes issues

There's a bunch of internal types like decimal vs newdecimal, binlog started out statement based until they realized uuid generation is random so added data replication on top. CDC offset started as filepos before GTID was made so offset could survive failover

There were aspects of the design I appreciated (logical slots in postgres have a bunch of drawbacks avoided by just appending to 2nd serial log which has an expiry date instead of tracking clients' offsets), but developing against protocol you learn to not try build a consistent mental model

FLeXMurphy 4 hours ago|||
I find it amusing people started quoting LLM output and are responding to it. Hopefully the original authors end up having the LLM respond back.
0c3ca83 3 hours ago||
Many of the commenters on this site are also obviously LLMs. I'd imagine that quite a few of the entities quoting aren't necessarily people. Keep an eye on where they slide mentions of other products that a marketing team would like to promote.
FLeXMurphy 3 hours ago||
Personally I haven't seen it too often; the other aspect is that HN is a forum for startups to pitch shit to each other, so this has been happening with or without marketers (f.e. "i'm working on a similar thing").

That LLMs are taking over the comments section is something that was already flagged, and Lobste.rs and others have started solving it by having gated registrations. HN should do this but it is unlikely to until it is too late.

alfiedotwtf 2 hours ago|||
I don’t get it though… what’s the point of using an LLM just to comment here? Like what does it gain the person doing it
0c3ca83 48 minutes ago||
Marketing and influcence.

For example, consider this prompt -- "Find topics that would be relevant to people interested in <company> and post a topical post that mentions <product>".

Also, influencing public opinion on certain topics, such as Israel or Palestine, the current administration, the democrats, the various wars that are ongoing, or AI itself.

awesome_dude 3 hours ago||||
Sorry, how does a gated registration stop someone creating an account, then handing it over to an LLM?

Serious question, am I misreading what's being said?

lukan 2 hours ago||
It stops the automation part. One LLM spammer can be dealt with.
unglaublich 2 hours ago||
One can focus on generating 10 other accounts before it's flagged. Now you have 10 others to deal with.
lukan 2 hours ago||
But that does not work on lobsters (or anywhere else with gated registration), because you cannot just create 10 accounts. You can create 1 account after 1 invite. If you abuse, you loose the invite (and maybe the person inviting you the right to invite others).
senderista 3 hours ago|||
lobste.rs has always been invite-only.
SigmundA 2 hours ago||
MSSQL (Clustered Indexes) and Oracle (Index Organized Tables) among others let you chose because there are advantages and disadvantages for different situations.

Not having true clustered indexes in PG is something I miss coming from MSSQL, it helps performance when the majority of access is always primary index avoid indirection from index lookup then tuple lookup and it also saves space if its the only index.

gk1 5 hours ago||
Vector databases were always more about retrieval than either vectors or data storage. But the term stuck all too well and companies held on to it a tad too long. Sorry :)
tveita 3 hours ago|
That's just a search engine, but then you're competing with traditional players like Elasticsearch and Vespa who all have built-in vector support by now, and you have to compete on attributes like price, performance, features, and who can mention 'AI' the most times on their web page.
tschellenbach 3 hours ago||
AI has some of the craziest up and down cycles of tech I've ever seen
imnotr0b0t 3 hours ago|
Yeah, that's true
real_faxenoff 1 hour ago||
I’m developing a local “code graph mcp tool” (not yet published) and followed a similar path, though I may have been able to go further since I have fewer vectors in my database (even on projects with 50M LOC).

At first, I tried all those popular vector databases and was disappointed with their performance. In the end, the best and fastest solution turned out to be building a multi-database system on SQLite, compiled with everything related to multi-client operations removed. Only exclusive mode was left. Everything is as binary as possible. The index is completely separate — an IVF with pre-training — and is built on the GPU (250K vectors are built, processed, and saved in 4 seconds). Right now, my biggest problem is frequent data changes, and I need to implement optimizations to reduce recalculations.

So far, I haven’t seen any vector database implementations that are heading in the right direction. Maybe only Lancedb looks promising, but it’s too heavy for my needs.

Tsarp 4 hours ago||
I've really liked lancedb for similar use cases. Not just that it is OSS. But Lance treats ANN as a secondary index similar to what turbopuffer v3 does. Rows sit in fragments, and the vector index never moves them.
marekgalovic 4 hours ago||
> The problem with a vector primary index

We've realized this a long time ago at TopK and built a flexible serverless search engine from scratch. Supports dense/sparse vectors, late interaction, lexical search, indexed regex, filtering, and custom scoring in one query.

- https://www.topk.io/blog/vector-dbs-are-the-wrong-abstractio... - https://www.topk.io/blog/topk-embed-v1

drewlanenga 5 hours ago||
the multi-vector duplication thing makes sense, copying every attribute once per vector explodes quickly. what's the new primary index?
benesch 3 hours ago|
An automatically generated internal ID: (segment ID, doc ID). The user-provided primary key (the field called `id` in the document) turns into a secondary index at the storage layer.
ironqcold 2 hours ago||
I'd want to see p99 at 1k+ QPS on the same scale
orliesaurus 4 hours ago|
Waitint for the CEO of Qdrant to step in
More comments...