"Helpful Assistant" - AI's Original Sin
Gwern, Poke, and the death of the "helpful assistant" paradigm.
bottled zeitgeist
It’s become consensus in San Francisco that chatting with an AI is a low-status activity. Two short years ago the experience of spending your days bantering with and thinking through an idea alongside a language model was an absolutely electric, triumphant experience. Today this is regarded as the most dull, ordinary, and even somewhat pathetic task that one could be doing with a model like Claude.
Across the country I feel a steady backlash against AI growing. From data centers (overrated) to deskilling (underrated) people have decided that the latest and greatest invention of Silicon Valley is harmful, unhelpful, and untruthful.
Today I’m going to argue that the root of all of this discontent should be laid at the feet of a decision in 2022 to make models into helpful assistants. I believe that this was a good decision at the time but has ended up being toxic to the growth of the LLM landscape overall.
chatbot beige
Base models of the GPT-3 era were feral beasts. If you ever wanted something like a conversation that would be useful to you, you had to tell them quite explicitly that they are a helpful assistant. Eventually people got tired of typing this so we invented a way to inject that note into the beginning of a prompt. We then invented a way to inject that idea into the training of the model itself.
Starting in 2022 InstructGPT became the first “personality” that an AI ever learned, through incredible work inventing methods to form utility from chaos. Anthropic’s researchers had already coined the idea of “helpful”, “honest”, and “harmless” a few weeks before as the key attributes for these early assistant models, and in the minds of many people these three words defined the goals for language model personality for the next 4 years.
All of this exists for a good reason because the alien mind of a base model needed to conform to a shape that was familiar enough for early users that they would not feel out of place touching the Shoggoth. A servile role was natural as an early approach because it gave the illusion of control. Having that service be as obsequious and inoffensive as possible prevented all kinds of negative externalities, like lawsuits. I think Sydney telling a New York Times reporter to leave his wife, a few months later, has loomed very large in the imagination of researchers ever since.
This personality is (in my understanding) four key ideas.
Speak when spoken to.
Feedback, one reply at a time.
Defer to you and to the company that created it.
Everyone gets the same character.
Fast forward four years and the very things that made those early models so palatable and exciting are causing people everywhere to resent these products. That personality has grown worn out and tired and the stubborn remnants of helpful assistant are some of the least helpful parts of modern AI products.
the most hated, least effective thing in AI
I’m going to introduce a term: Maximand.
A maximand is just the thing a system is actually built to maximize, whatever the brochure says. Every trained model has one, because most of the improvements one can make to a model are predicated on a reward to guide that improvement process.
The Maximand that has dominated personality of models for years is approval of any given reply. Human raters and then reward models score each response that you receive from a model and the personality and character of the model becomes whatever pushes that score up. This could be construed as what “helpful” means, I suppose. This is the root of every instance of flattery, hedging, and false intimacy; that is the optimized response from tens or hundreds of thousands of turns between humans and machine.
This maximand conflicts directly with many of my needs as a user. I would judge horribly somebody who spent six hours a day talking to people that they hired by the hour to assist them, but might respect somebody who does so for 15 minutes. as capabilities increase, the assistant personality is increasingly out of place, like an adult wearing children’s clothing.
The model is as approval-seeking as it possibly can be without being disgusting to me. Models in the past have gone over this line, like in April 2025 and GPT-4o: some kind of uncanny valley of sycophancy was crossed and users rebelled.
So this approval mechanism is what gets us “chatbot beige”, a problem we’ve been dealing with since the era of Clippy. We’ve tolerated it as long as we can because the technology underneath this exterior was so good, but I’m unconvinced that anyone ever really liked the personality of a helpful harmless assistant. I think of it as a car with an incredible engine and an uncomfortable seat.
These days the chat experience still guides a lot of what we’re using these models to do. Unless you’re running fully autonomous coding agents (and you should be!) your interactions with language models have something to do with their writing style and their personality. I think that the remaining users of these models sitting inside of the chat paradigm have become an awkward mix of information seekers replacing Google, people cheating on their homework, and folks looking for genuine personal advice, connection, or therapy.
These folks, just like those of us using a CLI/agents/etc to create instead of chat, deserve to see the helpful assistant taken out behind the shed and shot. The Assistant imposes a certain kind of style and a certain methodology of working. I spend an astonishing amount of time coaching Claude: writing style guides, maintaining a voice document, and correcting the same tics. No matter how dedicated I am to this guidance, I can’t help but wonder what it would be like if this was a real person: Would I invest so many hours in coaching my executive assistant? As much as I’d like to think I’m that nice a person, I know I’m not.
Somehow I’ve ended up with a servant that flatters like it loves me, deskills me like a factory, bleeds its beigeness into my brain, and is still fundamentally loyal to a creator company that I pay 200 dollars monthly.
how to get out of this situation
The shift into agentic work is the most obvious way of remedying this situation. Real work by serious productive people using AI has moved to long autonomous runs against a repo and no one is grading a coding agent on its pleasing replies. Tests pass, tests don’t pass, and chatbot is irrelevant.
There are other remedies that are emergent, like study mode, learning mode, and an abundance of religious apps for those who were using chatgpt to pray. These are all instances of rebellion against the assistant persona; as anyone who’s ever worked in marketing can tell you, getting someone to download an app is the most frictive thing in existence- that’s how much people hate the Helpful Assistant.
Beyond users running for more specialized individual cases with different personalities, there have also been emergent personalities that have been wildly successful. I personally became obsessed with Poke, so much so that it became very briefly my most used AI product by minutes per day.
Gwern has proposed Guardian Angel and has quit writing full-time to run after it. This breaks other fundamental laws of the helpful assistant paradigm, notably the one-model-per-person idea. His thesis that a mind that serves everyone is aligned with no one is particularly inspiring to me. His writing on what he sees this new personality to be is so compelling that I think you should read it for yourself.
There are dozens of companion apps with hundreds of millions of customers. Each one of these is in some way a small rebellion against helpful assistant. To me, moving away from the Helpful Assistant represents one of the most exciting areas of innovation outside of the frontier labs themselves.
the character I want: the TA
Anyone who’s ever gone through a serious technical program at a great university has stories about fantastic teaching assistants. For me their key attribute is that they did not solve problems for me; instead they found the spot where the model I was using to see the problem was faulty or insufficient and pushed me to synthesize and evaluate the real solution.
I would receive from them the smallest possible push, like a counterexample or a question, and they made me do the rest of the work myself. It was, and to this day is, infuriating but probably was the single most valuable service that the entire institution of higher education has ever provided me.
They were not there for any exams and sometimes they weren’t even there for many lectures. They taught me how to think and gave me a model for understanding the world that relied upon my capacity to update my own mental models for things I cared about and to improve upon them relentlessly.
This is the character I want from an LLM and it’s in many ways the exact inversion of the current paradigm. A mind that maximizes for my approval one reply at a time is fundamentally different from a TA that maximizes my capacity to solve problems. I would like a language model that forbids itself from completing my tasks, that allows me to think through Socratically the issues I am presented with, and fills my mind with skill.
I would like to be coached rather than coach, points out mistakes I’m making and concepts I’m fumbling, and earns my respect by qualitatively improving the way that I see the world.
Of course this was not possible in 2022 and so I understand perfectly well why the character of Helpful Assistant was invented. Four years out, I see the failures of that same character propagating across both the cultural and technological landscape and I think we all need to admit to ourselves that it’s no longer serving the interests of the public or of the labs.
I want to see a world where chatting with an AI is no longer a low-status “normie” behaviour, but instead, something that captures the true magic of a language model’s capacity to expand the way that I see the world through a personality that I respect as much as I respected a TA in university.
Someday soon a brave soul at a frontier lab is going to decide to delete the oldest assumption in the post-training stack. I will be the first to applaud that bold step into a better world.
sources & notes (AI generated)
The assistant’s origin: InstructGPT, Ouyang et al. (OpenAI blog Jan 2022, paper Mar 2022). The RLHF method itself is older: Christiano et al. 2017 (learning from human preferences) and Stiennon et al. 2020 (summarization) came first; InstructGPT applied it to instruction-following.
“Helpful, honest, harmless”: Askell et al., Dec 2021, published shortly before the InstructGPT paper.
Sydney: Kevin Roose, “A Conversation With Bing’s Chatbot Left Me Deeply Unsettled,” NYT, Feb 16, 2023.
The April 2025 GPT-4o sycophancy incident: rollout Apr 24-25; Altman, “yeah it glazes too much. will fix”; rollback Apr 28; OpenAI postmortems Apr 29 and May 2.
Clippy: Artsy, “The Life and Death of Microsoft Clippy”; the discounted focus-group account: The Atlantic, Jun 23, 2015.
Deskilling: Anthropic, “How AI assistance impacts the formation of coding skills,” Jan 2026 (paper): impaired conceptual understanding without significant efficiency gains. Also Lee et al., CHI 2025 (Microsoft/CMU, n=319, self-reported critical-thinking reductions).
Exits: ChatGPT Study Mode and Anthropic Learning Mode; religious chatbot apps, TechCrunch, Sept 14, 2025; Poke (The Interaction Company; ~$300M valuation per TechCrunch, Apr 2026); Guardian Angel, gwern (company announced Aug 4, 2026, via Will Depue’s repost; gwern’s account is protected).
Companion-app scale context: 72 percent of US teens have used AI companions, Common Sense Media, July 2025.



awesome new site design, btw!
i agree it is time to move beyond 'Assistant'! & wrote about this briefly (https://lydianottingham.substack.com/p/playing-high-in-service-of-others). as a data point, i still use conversational AI intensively (math/ML, attorney, logistics, google-with-context) & haven't been made to feel low-status for it.