

Join the discussion
Write your take first — we'll ask for email only when you're ready to publish.
- Hacker News
- > I asked Claude for its thoughts, and it doesn’t mince words: “what makes this genuinely interesting — and, frankly, a little embarrassing for the field — is that none of the ingredients are exotic.” The TL;DR is that someone just did a much more thorough job applying all of our known tools. In short: the sort of things that attack AIs are wonderful at.
did i just read two summaries/TLDRs (in a row) of the already-two-sentence summary right above?
- There a constructive way to leave feedback you knowby Ar-Curunir
- Does anybody have a filter for cryptography posts of the form:
AES IS BROKEN: Making giant assumption XYZ and requiring a less capable AES in ABC way, we've reduce the amount of operations needed to break AES from 10^X to 10^X-1!
These are tedious for people who are interested in cryptography but are not researchers in the field. (For the researchers, this type of thing may be useful.). The fact that AI is now "generating crypto results", suggests that soon we will soon have crypto-post-slop as clickbait...
by iansmith_hn - The prompt used to get this result is pretty crazy.by bawolff
- I also just finished writing a note - https://mkagenius.substack.com/p/notes-on-mythos-breaking-ae...by mkagenius
- >They [anthropic] appear to have just told it to get some results and then strapped its nose to the grindstone until it found some.
it is fun how well this works.
i cant find the link immediately (will look and edit with it), but somewhere in the "hello there the jacobian conjecture is false thanx" thread, someone brought up a different conjecture breakthrough where the prompts were basically just repeated "no, keep going" until a result was found.
edit: https://chatgpt.com/share/6a60b2eb-0b64-83ee-9c76-7931ca1de0...
i especially like "you should do a breakthrough". each prompt is less than ~20 words. makes me really question the whole "prompt engineering" stuff.
- Oh boy. The human here is effectively assuming the role of a Magic 8 Ball…
What a weird species of halting problem…
by alwa - Could send random characters along with 'keep going' and it would change nothing, the seed is whats causing deviation between these responses.by swordsith
- That second person stated that for many years they tried that particular graph problem on various AI models, starting with o1 and o3.
Its quite likely they now found the counterexample with a more serious prompt, and then for virality re-tried a few times with meme-prompts like "you should do a breakthrough", knowing that the model is capable of solving this particular one. Worst case the meme-prompts don't work and they share the real one they initially used.
by dist-epoch - Same with cybersec. They're finding bugs a human could've found if they looked hard but humans don't look hard at 100% of the code and the LLM can, at high speed.by inigyou
- In response to your edit, you should check out Terry tao's chat gpt logs about the recent Jacobian result. The models are smart enough to brute force some things, but can cut to the meat much faster with good promptingby adamzenith
- > i especially like "you should do a breakthrough". each prompt is less than ~20 words. makes me really question the whole "prompt engineering" stuff.
This is well past prompt engineering and into process engineering like six sigma. Just like in an early industrial revolution factory, we’re all still figuring out what works in the process of making stuff except this is so early that even simple things like “make this screw standardized” (or “no, keep going” in this case) is really high impact.
The degrees of freedom an LLM has is so large that we're going to be exploring their capabilities for decades, especially if they continue to get better. This is why IMO experts are always going to be better at LLMs in their field because they can force them LLM into processes (think prompt engineering -> CC dynamic workflows) that follow their work processes and get much better results out of them than “keep going.”
by throwup238 - > both outputs of Claude Mythos, their (still) unreleased advanced model
That sentence gives the impression that Mythos might be released in the future. That's clearly not going to happen - it's already "released" in as much as selected, trusted partners can access it, and the rest of us get it in the form of Fable - which is Mythos but with filters that downgrade you if you try to use it for anything even remotely related to cybersecurity or biology.
(The other day Fable 5 downgraded me to Opus after I asked it to explain the difference between tusks and teeth.)
by simonw - Genuine question here, why would you ask Fable to explain the difference between tusks and teeth? That's a task that can probably be handled by Haiku.by free_bip
- Got downgraded from Fable to Opus after asking about the Great Oxidation Event [1], presumably because it started thinking about cyanobacteria, which made the monitor afraid I was trying to make a biological weapon or something.by gwd
- This is good:
> If you’re under the impression that these models are “glorified autocomplete” or that progress is slowing down, I need to urge you: stop thinking that. The models are very intelligent and capable, they are getting better at a fast clip. I can cite measurable and impressive progress over just the past five months on specific types of problem I’ve asked them to look at. [...]
> On the other hand: if you think that models are super-intelligent or that AGI is already here, you should also stop thinking that. Working with these tools is like swimming in a pond where the ground drops off sharply. One minute you’re wading comfortably and there’s support under your feet. Then suddenly you cross a specific line, and you’re back to swimming on your own.
by simonw - > AGI is already here
I feel like there has been a ton of noise about this, but frankly, no one has actually defined what AGI means. I feel like the goal post is constantly shifting.
Take for example Humanity's Last Exam. It is so broad and complex that while an individual in a specific field might be able to answer their specific area of questions, they certainly would not be able to achieve >50% on the total question set.
There is this idea that AI has to be perfect to be intelligent - but we consider Humans intelligent and they are not even close. So is it the ability to generate novel ideas? Prove theorems? Pass tests?
I am not arguing that rote memorization is intelligence, or that we have achieved it, but does anyone know what AGI actually.. is?
by sph87 - I'm also getting irritated with the “glorified autocomplete” comments. Since nobody can post such comments and also use the tools I'm using, I'm wondering if the phenomenon is due to people only having experience with the free version of whatever it is they're trying to use?by dboreham
- There is grotesque hype for this new category of software. And the skittish AGI debate generates more noise than it's worth.
Do use the tool if it offers material/measurable improvement for your use-case, but budget carefully and keep your people. They probably do actually know what your material/measurable use-case really is.
If you only talk to yezbotz, you may zuck yourzelf into a corner and look like a zhmendrick.
- I feel like the development of AI has really shown what a gigantic spectrum intelligence actually is.by movpasd
- My mental model is this:
There is a vast ocean of human knowledge, far beyond the capacity of any human brain, even within specialised fields.
Books helped "plug the gaps" in our knowledge, increasing the scope that a single human mind can encompass.
Web search engines did the same thing, but more and faster.
LLMs are like search engines on steroids, essentially a research librarian that operates at 1,000x human speed and can "in context" locate relevant information, adapting it to fit the hole it needs to go into as well.
It feels less like discovering new theorems, but instead having direct access to all theorems, which is hugely valuable in itself.
I.e.: the recent counterexamples to open conjectures has largely been about the AIs "trawling through all the things" and scraping together every bit of human-generated knowledge ever produced that is relevant to the conjecture.
Conversely, in the past, we had to "make do" with sub-standard solutions where the problem had been solved, but finding every relevant solution in the ocean of knowledge was prohibitively time consuming.
In some sense, LLMs will "raise the floor" in what is considered the minimum level of quality of a solution, where even throwaway / toy designs will now start applying every bit of accumulated wisdom instead of just some of it.
We have mechanised attention.
by jiggawatts