Rendered at 22:04:00 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
janalsncm 23 hours ago [-]
> [in 2020/2021] the dominance of autoregression was not as well-established as it is today: GPT-3 had turned some heads, but the ‘ChatGPT moment’ wouldn’t come until late 2022
I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.
(The only exception I will make is encoder-decoder models which now are often done by decoder-only.)
But what made it go mainstream was RL. RLHF at first, then other improvements like DPO that were less of a pain in the ass to set up. Adding diffusion on top of that would be an even bigger pain in the ass.
Before ChatGPT there really wasn’t much of a concept of pre-training and post-training. It was all pre-training. Post training was what made the bots conversational and not just “continuing the thing you wrote to them”.
So in short, diffusion never took off because it was just a more complicated way to generate tokens, and the real problem was getting tokens in the right distribution.
matusp 17 hours ago [-]
> I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.
He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models.
> Before ChatGPT there really wasn’t much of a concept of pre-training and post-training.
Again, people spend years just post-training BERTs in various ways.
janalsncm 15 hours ago [-]
GPT2 and 3 work via autoregression. In chat bots, decoders work via autoregression.
> people spend years just post-training BERTs in various ways
Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.
Yoric 10 hours ago [-]
Asking out of curiosity, because I have limited experience in that domain.
I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?
janalsncm 2 hours ago [-]
It’s both. Fine-tuning a BERT model changes its weights, which causes the embeddings to change.
For example you might have one model which embeds a text query and another model which embeds an image. You also have a dataset of image + text captions. Training means updating the weights of those models so that the embedding of the image is close to the embedding of its corresponding caption.
benanne 9 hours ago [-]
I suppose what I was trying to say is that in the research community, people were still a lot more willing to entertain alternative modelling paradigms for language at that point, there was much less of a monoculture than there is today. It was really only after ChatGPT that non-autoregressive language modelling came to be seen as a fringe pursuit. Of course that is a highly subjective assessment, and my perspective is inevitably coloured by my own research interests at the time, and those of the people around me.
For what it's worth, I don't believe post-training diffusion models (of language or otherwise) is uniquely difficult, just a lot less thoroughly explored so far.
janalsncm 2 hours ago [-]
That’s fair, but my impression is that AI is a very industry-driven field. When commercialization fails, industry funding dries up and drags gov funding down too, resulting in an AI winter.
And industry really cares about solving practical problems. So as soon as autoregression “solved” the text generation problem (albeit imperfectly) the field was ready to move on. If text diffusion is not faster, cheaper, or higher quality, there is no ROI.
In fact text diffusion is much faster, so I have used it in the past. And I think it has the potential to be better as well, given a constant time budget. So I think it can be a really promising method moving forward.
echelon 23 hours ago [-]
> GPT2 was considered too dangerous to release
This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat.
Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.
They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.
onion2k 14 hours ago [-]
That's what's actually dangerous.
It's a different sort of danger. Knives, guns, etc are locally dangerous. The threat from AI is globally dangerous. Despite being a much lower risk of it happening, the blast radius is massively bigger, which means the overall danger is greater.
Also, maybe both things some be considered. Many countries manage the sale of knives, guns, etc. You don't need to accept that risk just because AI exists, or vice versa.
Nevermark 10 hours ago [-]
> Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.
You are hyperbolically judging a post-hoc cherry-picked subset of a great spectrum of concern, with hindsight experience nobody had at the time.
Safely navigating the future isn't an accuracy contest. Risk mitigation has to account for the distribution of costs for different signs of error, where novelty, uncertainty, and any potential for compounding effects all greatly multiply the need for hedging.
hodgehog11 22 hours ago [-]
I agree that this should be something that researchers reflect on. GPT-2 is one of the primary models to research on nowadays, and many recent developments have come from studying it as a test bench.
Imagine if CRISPR was considered "too dangerous" to publish because of the potential ethical ramifications, and that only a special few should be aware. It is utter self-righteousness, and it is shameful behaviour. The world cannot adjust itself to what it cannot see, so you risk greater catastrophe by keeping it secret.
The open dissemination of knowledge at every increment is the only way for society to truly deal with what is to come.
Yoric 10 hours ago [-]
As I understood it at the time, the rationale was that GPT2 made it too easy to produce content at scale and could easily be used for propaganda, phishing or other kinds of cons.
Which turned out to be true.
CamperBob2 5 hours ago [-]
And which says more about people being stupid and gullible than it does about AI models being smart or dangerous.
I'm tired of allowing stupid and gullible people to determine how the rest of us live, work, play, and create.
bandrami 13 hours ago [-]
I gave up on gun control as a political question when I realized you can buy a flamethrower at Home Depot
13 hours ago [-]
janalsncm 23 hours ago [-]
My impression at the time was they were perhaps overly cautious but this was a bunch of researchers who wanted to self-regulate. Anthropic didn’t exist, deepmind was also much less product-focused and relatively cautious. Chinese models weren’t really a factor either.
Government regulation was really not in the picture either in 2020 or 2021 tbh. The government was still trying to beat a pandemic. Some of the Biden admin eventually wanted to but it wasn’t very serious.
alightsoul 21 hours ago [-]
why are chinese models a consideration at all?
20 hours ago [-]
idiotsecant 20 hours ago [-]
Parent post is saying there was no community of peer models, just some people figuring it out as they went and with lots of slack to go slow if they wanted
alightsoul 19 hours ago [-]
Oh it's the typical china race thing, we can't let the antithesis of the us be better than the us etc. American exceptionalism
idiotsecant 19 hours ago [-]
The 'china' part is irrelevant. You're addicted to being mad. Parent post is just saying there were no other players
alightsoul 17 hours ago [-]
There were other players, china included, but it wasn't relevant until LLMs became profitable? China is being singled out as usual because it's the antithesis of the us
hansvm 16 hours ago [-]
But... there weren't. The other "players," China included, had duck all to show for it. Did they have the same potential capabilities? Sure, maybe. It wasn't just a profitability thing though. Their models sucked. China is being singled out in that comment because their models are exceptional, with very few competitors, so when you're talking about a history of the competition not being any good you _have_ to acknowledge one of the only good present-day models.
vikramkr 7 hours ago [-]
They were worried it would make it easy to generate a flood of misinformation. They were correct.
latentsea 20 hours ago [-]
What are your thoughts on the HuggingFace incident?
0xdeadbeefbabe 18 hours ago [-]
Optimal next token generation disguised as something more.
CamperBob2 6 hours ago [-]
What else is there?
CamperBob2 19 hours ago [-]
"Drama in search of a moat" is hard to beat.
slashdave 20 hours ago [-]
Well, no. Without controls, a language model can drive sensitive individuals to violence or suicide. The idea of releasing a frontier model without RL is frightening based on what we have learned.
anon373839 20 hours ago [-]
> The idea of releasing a frontier model without RL is frightening
In case you were not aware, strong base models (no post training at all) have been available for quite some time now. Including ones that eclipse “scary” frontier models from even a year ago.
SV_BubbleTime 19 hours ago [-]
So…
Words out of a magic box on a computer can’t make you kill yourself. Perhaps the people that would use that as encouragement are already mentally ill enough that it really doesn’t matter what the trigger is?
On a more callus but fully serious note, evolution starts out physical, that the species that don’t eat and breed as well, die. What happens when you remove that? When life is so safe that you basically wont stave, get eaten, catch a disease, don’t really need to compete all that hard for scarce resources, etc? Do you think evolution just stops? Or perhaps does a social or mental evolution become the predominant differentiator for successful reproduction over time?
p1esk 23 hours ago [-]
It’s refreshing to read something not AI generated.
benanne 8 hours ago [-]
Glad to hear it! My stubbornness about this makes me feel like a luddite sometimes, but the additional effort required is probably still worth it, for the time being.
2001zhaozhao 21 hours ago [-]
I would love to see models that can think at different rates and also output a thinking scratchpad alongside output text instead of before all output.
Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead.
17 hours ago [-]
vatsachak 22 hours ago [-]
I feel like there is still low hanging fruit on the auto regressive LLMs; the encoder
ainch 21 hours ago [-]
A great read - as with all of Sander's diffusion posts.
Marchant_hq 22 hours ago [-]
CDLMs sound promising for smoother, more coherent text generation. Excited to see how they tackle the token-level discontinuities.
ovin_dal 22 hours ago [-]
Diffusion models for language felt inevitable. Imagine the creative potential once these mature beyond current limits.
NickNaraghi 1 days ago [-]
I wonder if we’ll get something like CDLMs for automated harness engineering, sort of piloting the LLM underneath.
ramon156 23 hours ago [-]
how do tools like hermes do this? does it just review sessions and rewrite markdown files?
also haven't read too deep into the deepseek agent harness but the math in there was really cool. it sounded promising, at least.
latentsea 20 hours ago [-]
There's a cool research project https://github.com/exoharness/exo that is designed specifically as a self-modifying harness, so architecturally it separates things in a way that makes it a lot harder for the agent to break itself when modifying itself. It looks pretty neat. Saw the author do an interview on a podcast explaining it.
pests 20 hours ago [-]
> does it just review sessions and rewrite markdown files
Yes, same with openclaw etc. Some might have plugins to integrate with graph or vector databases besides only markdown files.
21 hours ago [-]
LoganDark 12 hours ago [-]
I love the idea of diffusion language models. I think they are potentially superior to autoregressive models, in the way that practically, considering whole systems tends to beat optimizing any one individual detail. To me, the way autoregressive models are sampled feels very fundamentally limited, and diffusion feels much more coherent by comparison.
nullbio 15 hours ago [-]
Conspiracy time: The continuous diffusion model has fruit. Frontier labs were trying to bury it with discrete diffusion distraction, because it is what they're using internally. Momentum wins, in the end.
ACCount37 11 hours ago [-]
No conspiracy. Diffusion is just a more finicky, more expensive way of generating the same tokens as autoregressive decoding.
amelius 1 days ago [-]
"Attention is all you need" should be renamed into "Attention is sufficient but not necessary".
p1esk 23 hours ago [-]
It’s the opposite: attention is necessary but not sufficient.
slashdave 20 hours ago [-]
Just about anyone building a large diffusion model is relying on transformers
ViscountPenguin 24 hours ago [-]
A quick look at the continuous diffusion models linked in the post shows lots of transformer models still
I disagree with this. Decoders were absolutely dominant in 2020 for chat. GPT2 was considered too dangerous to release, and I remember scrambling to get on the GPT3 waitlist. It worked.
(The only exception I will make is encoder-decoder models which now are often done by decoder-only.)
But what made it go mainstream was RL. RLHF at first, then other improvements like DPO that were less of a pain in the ass to set up. Adding diffusion on top of that would be an even bigger pain in the ass.
Before ChatGPT there really wasn’t much of a concept of pre-training and post-training. It was all pre-training. Post training was what made the bots conversational and not just “continuing the thing you wrote to them”.
So in short, diffusion never took off because it was just a more complicated way to generate tokens, and the real problem was getting tokens in the right distribution.
He's not talking about decoders, he's talking about auto-regression. Before ChatGPT, the dominant paradigm was fine-tuning BERT-like models.
> Before ChatGPT there really wasn’t much of a concept of pre-training and post-training.
Again, people spend years just post-training BERTs in various ways.
> people spend years just post-training BERTs in various ways
Yes, I was one of them. That’s not called “post-training” it’s called fine-tuning.
I thought that fine-tuning was changing the weights in the model, not the embedding? Or did I misunderstand?
For example you might have one model which embeds a text query and another model which embeds an image. You also have a dataset of image + text captions. Training means updating the weights of those models so that the embedding of the image is close to the embedding of its corresponding caption.
For what it's worth, I don't believe post-training diffusion models (of language or otherwise) is uniquely difficult, just a lot less thoroughly explored so far.
And industry really cares about solving practical problems. So as soon as autoregression “solved” the text generation problem (albeit imperfectly) the field was ready to move on. If text diffusion is not faster, cheaper, or higher quality, there is no ROI.
In fact text diffusion is much faster, so I have used it in the past. And I think it has the potential to be better as well, given a constant time budget. So I think it can be a really promising method moving forward.
This is how ridiculous this industry is. Regulation-seeking panic over nothing. Drama in search of a moat.
Everything is "too dangerous". GPT2 is going to invent a time machine and break crypto and genetically engineer super rabies.
They sell knives, guns, combustible materials, and multi-ton heavy machinery in stores. That's what's actually dangerous.
It's a different sort of danger. Knives, guns, etc are locally dangerous. The threat from AI is globally dangerous. Despite being a much lower risk of it happening, the blast radius is massively bigger, which means the overall danger is greater.
Also, maybe both things some be considered. Many countries manage the sale of knives, guns, etc. You don't need to accept that risk just because AI exists, or vice versa.
You are hyperbolically judging a post-hoc cherry-picked subset of a great spectrum of concern, with hindsight experience nobody had at the time.
Safely navigating the future isn't an accuracy contest. Risk mitigation has to account for the distribution of costs for different signs of error, where novelty, uncertainty, and any potential for compounding effects all greatly multiply the need for hedging.
Imagine if CRISPR was considered "too dangerous" to publish because of the potential ethical ramifications, and that only a special few should be aware. It is utter self-righteousness, and it is shameful behaviour. The world cannot adjust itself to what it cannot see, so you risk greater catastrophe by keeping it secret.
The open dissemination of knowledge at every increment is the only way for society to truly deal with what is to come.
Which turned out to be true.
I'm tired of allowing stupid and gullible people to determine how the rest of us live, work, play, and create.
Government regulation was really not in the picture either in 2020 or 2021 tbh. The government was still trying to beat a pandemic. Some of the Biden admin eventually wanted to but it wasn’t very serious.
In case you were not aware, strong base models (no post training at all) have been available for quite some time now. Including ones that eclipse “scary” frontier models from even a year ago.
Words out of a magic box on a computer can’t make you kill yourself. Perhaps the people that would use that as encouragement are already mentally ill enough that it really doesn’t matter what the trigger is?
On a more callus but fully serious note, evolution starts out physical, that the species that don’t eat and breed as well, die. What happens when you remove that? When life is so safe that you basically wont stave, get eaten, catch a disease, don’t really need to compete all that hard for scarce resources, etc? Do you think evolution just stops? Or perhaps does a social or mental evolution become the predominant differentiator for successful reproduction over time?
Right now models need to rely on less legible compressed CoT to get high intelligence per token/step, but with diffusion they would just need to output more tokens per step instead.
also haven't read too deep into the deepseek agent harness but the math in there was really cool. it sounded promising, at least.
Yes, same with openclaw etc. Some might have plugins to integrate with graph or vector databases besides only markdown files.