Top
Best
New

Posted by tosh 11 hours ago

Web Search API(developers.cloudflare.com)
441 points | 205 commentspage 2
sreekanth850 9 hours ago|
Create bot detection and bot protection, then sell crawlers. Is this the peak of hypocrisy?
doginasuit 8 hours ago||
Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.
madibo3156 8 hours ago|||
However, so-called AI crawlers are not the same as crawlers of yore. They hit live pages every time a user prompt triggers a web search.

This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.

senko 4 hours ago||||
A crawler that respects robots.txt is useless in practice since many sites only allow Googlebot and maaaybe Bing - by name.
sreekanth850 8 hours ago|||
And you think all this web AI crawlers will respect robots.txt. That era is gone.
Jskewel 3 hours ago||
The big USA ones do, and it would be madness for them to do otherwise.

But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.

dawidpotocki 3 hours ago||
That doesn't match my experience. For example, I recently had Meta scrape robots.txt disallowed paths. On some page they claimed they respect it but they don't. And yes, all the IPs they used to connect to me (they used a different /64 network for each connection) were owned by Meta, it was not someone pretending to be them.

Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565

I ended up just fully closing connections with no response from these assholes on any URL.

0xbadcafebee 6 hours ago||
That's not peak hypocrisy, that's peak capitalism
yellow_lead 4 hours ago||
Reselling APIs seems lazy to me. I wonder if CF plans to make their own provider. That's what I had assumed when I read the title.
JMiao 4 hours ago|
why must it always be hard
freakynit 10 hours ago||
Tried one query on ceramic.ai (the default provider for cloudflare web search api): "qwen-3.8 flash next and rtx 5090 best inference setup" ... 0 results ... same query on google and ddg both yield proper results.

Then shortened the query to just "qwen-3.8 flash next" ... results came.. all unrelated. In fact, these were almost all paper links .... no relation to actual search term.

And I had thought that I finally had found a cheaper search alternative.

freakynit 10 hours ago||
You know the craziest part? This time I searched for their own website address: "ceramic.ai" .. results came... none pointing to the website or any page on it.

Then searched for "Cloudflare OHTTP Gateway" .. this text is literally in the title ... but zero link for this page.. the closest it yielded was this link: "https://developers.cloudflare.com/privacy-gateway/" ... it seems cloudflare updated this 2 days back.. the original content was last updated in 2022 ... so that's what the cutoff index seems to be.

scosman 10 hours ago||
I made a zero ads SERP using one of these "AI first" search providers: https://github.com/scosman/froogle (live version https://froogle.fyi). In this case Keenable.ai. Generally the same pattern: it's not usable.
hrideshmg 3 hours ago||
Surprised no one in this thread has mentioned running a self hosted search API.

I personally used to use Firecrawl's paid credits (got a bunch of em for free at an event) before I realized that they allow you to self-host your own instance (albeit missing some features I never use anyways).

It's been working really well for my agents, I even hosted a small observability tool that proxies the requests so I can see how many are failing and the percentages are always below 2%.

jeromechoo 2 hours ago|
Basically a self-hosted Google SERP scraper?
laumars 4 hours ago||
Lately I've been using SearXNG for personal models. It's free and seems ok thus far.

https://docs.searxng.org/

tom1337 11 hours ago||
I wonder if the three search engines get access to cloudflare protected sites without any captcha or bot interventions
astonex 11 hours ago|
Most likely not. Their Crawling service for example does not bypass the cloudflare protections either.
weird-eye-issue 10 hours ago||
You are conflating a couple of different things here

There actually is such a thing as verified bots on Cloudflare that gets through most blocks (and these services are likely are part of that), but ultimately it just depends on how the website owner has things set up in Cloudflare

viraptor 9 hours ago|||
> There actually is such a thing as verified bots on Cloudflare that gets through most blocks

Verified bot is just a label. What you do with that information is entirely up to you as the operator. It doesn't say anything anything about the service and doesn't provide any guarantees about the traffic.

weird-eye-issue 9 hours ago||
No it's not just a label because it's a category that gets used in the firewall settings. A lot of sites allow only verified bots and block non-verified ones. So yes as I already mentioned website owners can of course block verified bots but it's much much less common than blocking non-verified ones because most site owners don't want to accidentally deindex their site from Google...
viraptor 2 hours ago||
You get to decide per bot. Or per category: https://developers.cloudflare.com/bots/concepts/bot/verified... You can make for that Google or all indexers are always allowed.
dbbk 10 hours ago|||
I hate to be rude but once again I am begging people to actually read the post. It is mentioned in the second paragraph that they are all verified bots.
weird-eye-issue 9 hours ago||
Actually it doesn't say they are verified bots but it does say that they follow the verified bots standards. But yes I think we can assume they are verified bots which I already said in my comment so I'm not really sure what the purpose of your reply is?
EcommerceFlow 4 hours ago||
Spent a few months building a product scraper using a mad mash up of various LLMs, OCR, etc. The pricing for their providers is 3x-8x higher than something like Luna 5.6 w/ Web Search. Not sure what their differentiator is, unless they just wanted to launch something.
0fflineuser 8 hours ago||
I am pretty sure exa specifically say it trains on your data in it's privacy policy, so how can it be ZDR ?

I remember as I was looking at the available web tools for hermes agent not to long ago and looked through the keyless web providers privacy policies, which exa is one of them.

ashley95 8 hours ago||
An important enough customer can get special contract terms.
lukewarm707 7 hours ago||
it is not zdr via cloudflare, just a typo in the documentation. it is stated no zdr elsewhere on the page.
solaire_oa 3 hours ago|
Anthropic plausibly uses Brave Search... and Brave search maintains its own index. Makes sense: cheaper search API, leveraging non-Google, etc.

Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).

So what is CF providing here? Maybe some free credits to entice us to use their router? No, not that either ("billed to your AI Gateway credits"). Maybe a comparison of which agent search yields the best results? Nope.

It's a crappy proxy- probably less efficient and more volatile than hitting the agent API directly.

This is only if I understand the product correctly (which I admittedly skimmed) due to the sheer number of screeching vibey nothingburgers coming out of CF over the past month.

alexey-salmin 2 hours ago|
>Here we are, one layer of indirection more: Ceramic, Exa, Linkup. Who knows what they use. If you told me that those 3 build and maintain their own index, I'd first question whether that was true, and if it is, I would question whether it was any good (relative to Google/Bing/Brave).

Hi! Exa Head of Index here. We certainly do have our own index and it's one of the biggest among the independent players (i.e. not Google and Bing, which by the way closed off their official search APIs). [1]

Regarding the quality: search is a multi-dimensional problem, you can be better on one set of queries and worse on the other. There are tons of benchmarks in the industry, all the players in the AI search market are fighting very hard to climb to the top, updates are shared every week.

We track dozens of use cases and run evals continuously, we perform well on all the verticals we optimize for. Not only we top the ranking on e.g. financial queries, but also Claude with Exa search performs better that Claude with native search -- as measured by independent observers [2]. This means that the underlying search is materially better for the outcome, it's not just how we evaluate the search itself.

[1] https://lnkd.in/p/e9u3dyEG [2] https://lnkd.in/p/enYe4h7u

BrendanEich 50 minutes ago||
Subtweeting feels low, and I was not picking on you only. Here is a direct link: https://x.com/BrendanEich/status/2107224138394095965
More comments...