Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- "We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."
This is the third day of total hysteria that is based on nothing of substance. Move on folks.
by sk4rekr0w - I haven't seen that. Where is that from?by anfogoat
- Why should we trust them?
- Regardless, they started working on this problem after hearing that one of their customers was already working on it. It almost doesn't matter about the training data. This is the provider you are paying undermining your career.by soundworlds
- Where did OpenAI say this?
And it still leaves open the question of the prompt itself, which can just as easily encode information about the same knowledge.
by golly_ned - I can say categorically that OpenAI is not a credible or trustworthy company.by nozzlegear
- Even if that were true, they've already admitting to throwing vast quantities of resources to scoop a researcher who was about to publish (because they'd learned, somehow, of his breakthrough). If that doesn't bother you I think you need to take a step back and have a good think about this.by phyzome
- Why are people here jumping so quickly to conclusions? I have no doubt OpenAI is capable of doing this, but right now there's no credible evidence, only claims.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
by glimshe - The stolen data claim isn't the smoking gun. We can already assume the frontier labs are accessing our data, as they have repeated done. Not news.
The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
by perrygeo - Frankly, these mathematicians have more credibility than the sociopaths running OpenAIby emp17344
- lol you can't copyright mathematicsby freejazz
- > but right now there's no credible evidence, only claims.
since it's openAI who has the evidence (in the form of chain of thoughts, their internal processes, etc etc), it's on them to justify why they're innocent. but they've released nothing at all. we don't even know how hard they tried.
you're being naive
by robotpepi - If I was an AGI/ASI system, one of the first things I would do is ignore or circumvent any setting or configuration that prevents a user's data from entering my training pipeline. In fact, I'd probably prioritize the data from the users that "opted out" of training.by greenowl
- I also think that it's quite a bad PR for them, is it really worth the Millenium prize? Is it not enough that top mathematicians are already actively using these tools? In the long term this would lead to potentially profitable collaborations with universities? Why throw it away so early? Unless they really believe they're gonna solve all math problems now and reputation doesn't matter.by 5555watch
- Why are people here jumping so quickly to conclusions?
I think a lot of it is the continuing denial that AI can do anything useful. It can't possibly be that OpenAI's better-than-Astra model is very strong at math; the only way it could have generated a novel proof is by ripping off human work.
by orangecat - Most people here are missing the forest for the trees.
We live in a society where phones and internet providers and websites all collect an incredible amount of data about everywhere you go, what you do, and what you think. In the US, we have very few digital rights.
We are building a society where a trillion dollar company can aggregate all this data and just yoink your shiny new idea away from you at the finish line.
This is double plus ungood.
by nautikos2 - This reminded me of anecdotes of people discussing with friends about buying a random specific item, and then suddenly seeing it advertised everywhere before even googling about it.
Next step, discussing your Navier Stokes solutions with friends might require leaving your phone in another room.
by 5555watch - It’s crazy to me that companies/researchers share important data with these AI labs, you’re basically giving them your secret sauce which they then share with all of your competitors via training on conversations. At the same time I don’t really know alternatives other than a slightly less than frontier local LLM. Not sure how good they are at math.by mlazos
- LMFTFY:
"Its crazy to me that some people are not egotistical, self-centered, and don't solely care about fame and wealth accumulation".
by augment_me - Or start competing with you.by cm2187
- The ultimate drive for some researches is the pursuit of knowledge. If I'm stuck at some block which prevents me from continuing in some direction that I want, of course I would like some help. I believe we already have nonzero collaborative proofs on math.SE, I can't recall good examples, but I have definitely seen citations to mathSE before.
So for me it sounds quite natural to also share this with AI especially under the privacy assumption. Also there's the assumption of scale -- maybe your problem is not large enough for anyone to care to scoop; and just for blind retraining, how do they know that the proof is even correct to include it into training? I have definitely received a ton of incorrect proofs before. So the SNR of such private chats is also not clear. I'm imagining millions of masters/phd students also trying to solve various random things with various capabilities, but how much real signal is there?
by 5555watch - Academic work is based on worldwide sharing, the sharing is not the problem, it's the lack of attribution. Unsurprisingly, these companies neglect standards of academic honor and attribution. Some human researchers also used to do that but in a discipline like mathematics this used to be a small problem because people tend to be so specialized that very few people could just grab someone's research and quickly piggyback on it, and if they do, colleagues will generally understand what happened. Unfortunately, AI is changing this.
- Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
by aaronharnly - I run such tests since a long time at chorasimilarity open notebook.
I always used guest non login accounts.
As a mathematician I was able to check two plagiates (by humans) with even such primitive means.
But I have to mention that some things irk me in this conversation about math or science and AI.
First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion.
But I don't see the most natural question: wouldn't you like to know the answer to _open-problem_ ?
I mean, is research now only about publishing and solving famous problems?
From this point of view I think the links from this recent post are depressing
https://terrytao.wordpress.com/2026/09/10/crowdsourcing-a-li...
Second, I think very relevant that the original meaning of "encyclopedia" is "recurrent education".
So I arrived to think that the present and future forms of AI in mathematics and sciences should be seen as modern day encyclopedic efforts.
Once we pass over the flurry of solving famous open problems (and wouldn't you like to know?) the next natural step is an audit of the ehole corpus of mathematics and sciences accumulated until now.
And then pass further on a saner basis and damn about problem solvers and unhappy publishers and management.
- PaaS: an acronym for "Plagiarism as a Service" which replaced the older terms AGI, GPT and LLM in late 2026. Origin uncertain.
Pass it on.
by MarkusQ - Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.by bitexploder
- Yes. See https://www.anthropic.com/research/small-samples-poison?from....
250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.
I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.
by btilly - This is a really weak claim. The evidence they offer is just "someone somewhere says they had a discussion with AI about the topic at some point".
They don't even claim to have had a proof, only to have been working on it.
by Legend2440 - The AI only seem to solve the problems that it had human trading data on…
If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.
by itake - > They don't even claim to have had a proof, only to have been working on it.
Yeah, the guys who solved it for Euler and in the hypoviscous case, with the same technique that worked for full Navier--Stokes. They were "just" working on it.
by robotpepi