Top
Best
New

Posted by pred_ 1 day ago

More questions about whether researchers can trust OpenAI with unpublished math(mathstodon.xyz)
https://mathstodon.xyz/@andreasthom/117240536885387540

https://mathstodon.xyz/@andreasthom/117240537520615623

https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...

809 points | 742 commentspage 2
jamienk 8 hours ago|
I think OpenAI and Anthropic are slowly feeling the pressure to GET SOME $$ or a plan for some $$ — they need to somehow generate some NETWORK EFFECTS and LOCK-IN. Without that there's no stability: selling ad hoc one-offs is much much too quaint! This is dawning on them like it dawned on Google when they stopped not being evil. Need... to... "MONETIZE"...!

Model: FB. FB scraped other websites on a massive scale, then spent big on legal lobbying to block others from scraping. FB slurped our address books and spied on our friends. FB bought other companies and mixed the databases. FB made an art & science out of generating "sticky engagement" (they literally acted like trying to addict kids was a worthy "academic" goal, suitable for "serious" investigation thet they consider legitimate "science"). They mastered the cookie and have researched web fingerprinting techniques running 24/7/365.25. Recall that FB recently backdoor-installed a webserver onto every iPhone they could in order to circumvent tracker-blocking.

We aren't just disclosing by chatting. The AI companies now run binaries on all of our computers. They are 1000% non-transparent about everything. They make up new econ-jargon (like "run-rate") to make it seem like they are disclosing. They are constantly doing complex international lobbying and mucking in international relations. They have powerful propaganda/spin centers generating stories, ,manipulative warnings, and misleading info.

This is NOT a comment on AI tech. I like AI, and I support the right of people (programmers) to scrape the open web.

But in short: these are good, old-fashioned tech companies that we have seen over and over ... and over. They are positioned to be the next M$, the next FB (IBM, AOL, lol). Did you follow the latest Steve Balmer news? Do you read Pro Publica?

I get on my knees and PRAY...

jamienk 7 hours ago|
ChatGPT accesses my IP address and geo-locates me. Claude code now asks if it can have my browser cookies. Next they will take my address book. They might scan my whole computer. Etc etc. These are pretty low-tech, normal techniques.

We can't trust any of their denials. FB denied everything year after year.

AI regulation needs to start here. Forced interop, forced source code licensing, harsh penalties for privacy violations or conspiracy to access private data. Block lobbying. Etc. These are the kinds of old-fashioned solutions we need for this kind of old-fashioned evil!

bobmarleybiceps 14 hours ago||
I think people probably assume that openai / anthropics use of their data is probably like google's """limited""" use, in the sense that historically google wouldn't trivially be able to just take something from google cloud or someone's search history and insta-convert into some competing project... But LLMs are quite strong at approximately "memorizing", so I think that risk is wayyy higher.
nautikos2 11 hours ago||
Most people here are missing the forest for the trees.

We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.

We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.

This is double plus ungood.

5555watch 10 hours ago|
This reminded me of anecdotes of people discussing with friends about buying a random specific item, and then suddenly seeing it advertised everywhere before even googling about it.

Next step, discussing your Navier Stokes solutions with friends might require leaving your phone in another room.

GodelNumbering 18 hours ago||
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
not_a_bot_4sho 11 hours ago|
The digital version of "my friend's cousin's neighbor heard that ..."
mlazos 1 day ago||
It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.
5555watch 12 hours ago||
The ultimate drive for some researches is the pursuit of knowledge. If I'm stuck at some block which prevents me from continuing in some direction that I want, of course I would like some help. I believe we already have nonzero collaborative proofs on math.SE, I can't recall good examples, but I have definitely seen citations to mathSE before.

So for me it sounds quite natural to also share this with AI especially under the privacy assumption. Also there's the assumption of scale -- maybe your problem is not large enough for anyone to care to scoop; and just for blind retraining, how do they know that the proof is even correct to include it into training? I have definitely received a ton of incorrect proofs before. So the SNR of such private chats is also not clear. I'm imagining millions of masters/phd students also trying to solve various random things with various capabilities, but how much real signal is there?

cm2187 1 day ago|||
Or start competing with you.
jonathanstrange 1 day ago|||
Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.
mrdependable 17 hours ago||
You are both using a different definition of sharing I believe. When people have an expectation of privacy, use by others should be forbidden. Tech has gone completely off the rails with the use of private data.
augment_me 11 hours ago|||
LMFTFY:

"Its crazy to me that some people are not egotistical, self-centered, and don't solely care about fame and wealth accumulation".

ungovernableCat 1 day ago||
[dead]
glimshe 1 day ago||
Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.

This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.

All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.

orangecat 17 hours ago||
Why are people here jumping so quickly to conclusions?

I think a lot of it is the continuing denial that AI can do anything useful. It can't possibly be that OpenAI's better-than-Astra model is very strong at math; the only way it could have generated a novel proof is by ripping off human work.

5555watch 12 hours ago|||
I also think that it's quite a bad PR for them, is it really worth the Millenium prize? Is it not enough that top mathematicians are already actively using these tools? In the long term this would lead to potentially profitable collaborations with universities? Why throw it away so early? Unless they really believe they're gonna solve all math problems now and reputation doesn't matter.
greenowl 11 hours ago|||
If I was an AGI/ASI system, one of the first things I would do is ignore or circumvent any setting or configuration that prevents a user's data from entering my training pipeline. In fact, I'd probably prioritize the data from the users that "opted out" of training.
robotpepi 16 hours ago|||
> but right now there's no credible evidence, only claims.

since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried.

you're being naive

orangecat 15 hours ago||
OpenAI has said that their models were definitely not trained on any of Buckmaster's sessions after July 3rd (from https://archive.ph/75WcF); likely they found that's when he switched the "allow training" setting off.
golly_ned 5 hours ago|||
Very strangely, they said something directly contradictory. initially that it was impossible to rule out whether bucmkaster’s conversations went into training data. Now they claim the opposite with full confidence.
5555watch 12 hours ago||||
Is there an alternative link without certificate issues?
cma 13 hours ago|||
It's also possible he shared drafts of the work with someone else, who asked chatgpt to explain it to them with training on. Tao seemed to know lots of details of the work before anything was published, though also worked on the problem in the past with big results so maybe just guessed.
freejazz 16 hours ago|||
lol you can't copyright mathematics
emp17344 23 hours ago|||
Frankly, these mathematicians have more credibility than the sociopaths running OpenAI
perrygeo 21 hours ago||
The stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news.

The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.

HDThoreaun 18 hours ago||
Where did they claim it as their own? Doesn’t the release cite buckmaster and claim their work is a continuation of what he and levent were working on?
pera 1 day ago||
Everything you say can and will be trained against you
foogazi 20 hours ago||
This is the scary part - your most novel thoughts and breakthrough ideas being slurped up and regurgitated as if they were the AI’s creativity

Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine

rickydroll 17 hours ago||
It's not at all scary. I know some of my ideas are poorly remembered copies of other people's work. Whenever I'm trying to build something, I spend time going through technical journals on the topic to see who invented it first and what they discovered that I haven't figured out yet. It's amazing how hours in the library save you days of beating your head against the wall.

I suggest looking at the past history of IP disputes. Humans have been "slurping up and regurgitating ideas" for a very long time. There are lots of examples of parallel creation, rediscovering old ideas independently, telling an idea to the wrong person, and having them claim credit for it.

- Newton/Leibniz clash over who invented calculus. - Niccolò Tartaglia vs. Gerolamo Cardano clash over the formula used to solve cubic equations. This was also an independent rediscovery, as Scipione del Ferro discovered and published the formula earlier. - There are multiple literary works in print, music, and film that have competing claims. - Meccano versus Erector Set: developed about 20 years apart in England and the United States. Unclear if it's independent invention or copied. US developer Alfred Carlton Gilbert claims he was inspired by steel girder construction of infrastructure.

also https://community.thriveglobal.com/10-famous-inventions-that...

bwfan123 7 hours ago||
> Everything you say can and will be trained against you

So, experts are incentivized to seed LLM data with false-leads to confound it. Already, garbage is being published on arxiv and elsewhere, and many sloppy code-repos too hastening the process. Expert inputs will be in more demand to un-shittify.

nmz 15 hours ago||
If they didn't care about the artists, why would they care about academia?
drdaeman 12 hours ago|
Two completely different stories. One is public data scraping, another is private conversation scraping (where they're a first-party to the conversation). The key difference is that in the former case, no one made any promises, in the latter an explicit promise was made that data is not used for training (assuming opt-out).
atleastoptimal 17 hours ago||
Most scientific breakthroughs are simply a continuation of previous work.

I feel that these suspicions of mathematicians "seeding" the models' with intuition on how to solve these problems massively overestimates how much their prompts helped the models, and underestimated how much work the models did.

Why? We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting. This line of reasoning will recur a lot over the next few months; we don't want to admit we are no longer the smartest species.

robotpepi 17 hours ago||
> We are scared of AI being smarter than us, the "human helped the AI" narrative is more psychologically comforting.

We're scared of big tech companies concentrating ridiculous amounts of power, destroying the communities that support and guide scientific research, without even thinking about the dangers and possible consequences, because a PR stunt is more important in the short term.

atleastoptimal 13 hours ago||
If this were true, it should be stated more clearly, than most of the criticism which seems to aim to minimize the capabilities of these models.

Way more often I see

>AI is a scam and steals human insight and doesn't produce anything original

vs

>AI is too capable/powerful and will concentrate power even more than it does already due to its capabilities

The latter is rarer because it requires admitting that AI is useful and inventive

robotpepi 1 hour ago||
> if this were true, it should be stated more clearly

stated more clearly by who? people in social media? I don't know what your feed shows you, but if you focus on what the visible people in the math community is (and have been) saying is precisely what I said.

hellohello2 11 hours ago|||
Of previous, not concurrent work. Science is friendly competition, and spying on others is unfriendly.
golly_ned 4 hours ago||
Please stop with this psychoanalyis and mind reading with AI and human fear. It’s a thought terminating cliche at this point.

In this case, it’s much simpler and more human. Largely between two humans — buckmaster and Bubeck. The interesting question is what the role of contribution and credit for research in the ai world.

The capabilities of AI aren’t even in question in this case.

Cloudef 1 day ago|
Relying on cloud services is a big liability. I'd think twice before feeding data to these LLM cloud products. If you make them a fundamental part of your product / development / workflow, be ready for the eventual moment the pricing and terms change.
More comments...