It's also pretty cheap if you're paying for it (for coding). But I don't quite trust it for anything complicated, so I often use Terra.
redox99•Aug 6, 2026
Terra is awful and not cheap enough to make up for it. I'd suggest using luna or sol.
Sammi•Aug 6, 2026
So Luna is more awful but it's OK because it's cheaper?
redox99•Aug 6, 2026
Yes. You're better off running Luna Max or Sol Medium and never touching Terra.
LaurensBER•Aug 6, 2026
I've had great success with using Luna and having DeepSeek 4 Flash check Luna's output. Oh My Pi has a mode built-in that does this automatically ("advisor" mode). Deepseek only interrupts when it spots an issue so it doesn't slow Luna down.
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
dannyw•Aug 7, 2026
Yes, in the same way a basic, dry hotdog at $5 is bad, but I could love it in $0.50.
Pricing absolutely matters.
stuartq•Aug 6, 2026
Terra is far from awful. I've not used Luna enough to judge whether it's significantly better than Luna, but at least via GitHub Copilot, Terra is leaps ahead of Sonnet 5.
causal•Aug 6, 2026
I guess I should give it another shot because I had pretty bad experience with Luna when it first came out. Stuff Sonnet knew better.
timpera•Aug 6, 2026
Any improvement to the ChatGPT free plan is really nice. It's easy to forget that most people have never used a SOTA model, and only think of their experience with GPT-4o or Google Search's AI Overviews when asked about AI.
msq22•Aug 6, 2026
Were free users subject to some limits? I've never run into any limits.
timpera•Aug 6, 2026
It used to be 10 messages every 5 hours using GPT-5, then unlimited 4o-mini. More recently, it went down to ~5 messages per day on 5.5 Instant, then unlimited on 5.5 mini.
jauntywundrkind•Aug 6, 2026
At first I thought the Sol updates was perhaps trying to help with some complaints of Sol burning through tokens, complaints that have prompted some new data points on https://codex-resets.com/ .
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
laweijfmvo•Aug 6, 2026
What a bizarre example of a more direct answer. Any human would simply say “No, it doesn’t rain here in summer.”
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
taikahessu•Aug 6, 2026
Would you like to know more? Seriously though, of course, it's tuned for maximum engagement, not maximum efficiency. I wonder how long this engagement dopamine circus can last... too long apparently.
sunaookami•Aug 6, 2026
GPT models were RLHF'd to death, they will never give a final, direct answer. Every release since the GPT-4o catastrophy is like this, it's so tiresome. They need a complete reset before it can become actually usable again. Or not as maximum engagement seems to be their goal.
colingauvin•Aug 6, 2026
They must really be feeling the commoditization pressure. I'm not sure what the way out of this is, ChatGPT and Claude are still good products, but they are not necessarily premium products anymore.
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
redox99•Aug 6, 2026
I'm not sure I agree
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
colingauvin•Aug 6, 2026
>1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
davidguetta•Aug 6, 2026
especially when 99% tasks don't need fable, and certainly not the next model
user43928•Aug 6, 2026
I also scoffed at the ChatGPT $200 Pro plan back then.
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
thejazzman•Aug 6, 2026
Don’t say the last part out loud or it will be $800.
ignoramous•Aug 6, 2026
Every week, 1 billion people turn to ChatGPT for everything from quick questions and web searches to planning, research, advice, and complex decisions.
Guess, Google's AI Mode is chipping away at their consumers (I know I haven't used Chat in a long, long while for 'quick questions and web searches' after OpenAI did away with "think" which I always use). The money-minting office & coding market Anthropic has cornered is hyper-competitive at both the frontier & low-cost ends. OpenAI is reactive [0] and seems right up against it, despite the strength of its excellent models.
[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
skybrian•Aug 6, 2026
$20/month also gives you API access for coding, in any coding agent. I keep hitting the weekly limit but it's a good deal while it lasts.
porridgeraisin•Aug 6, 2026
Did they do away with think? I think now you have to do it with /think
drivebyhooting•Aug 6, 2026
Google’s AI has been very glitchy for me lately. I used to reserve chatGPT for serious work and Gemini for daily personalized unimportant things. But now I switched completely to ChatGPT and resigned myself to their memories/personalization.
sk4rekr0w•Aug 7, 2026
You've jumped to a naive conclusion
firasd•Aug 6, 2026
I think it’s a misread to think the default ChatGPT model switching to GPT 5.6 Luna is some sort of desperation move. Keep in mind that Claude .ai never had this extreme stratification between the frontier models and the free tier (Sonnet is available to free users with rate limits).
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
freakynit•Aug 7, 2026
Claude's rate limits are absolutely shit for free users. 3 chats and its gone for 24 hours.
ElijahLynn•Aug 6, 2026
I can't wait to never see a reasoning button ever again. Why do I have to reason about what reasoning level to use?
timpera•Aug 6, 2026
I think it's nice to be able to make the model reason for dozens of minutes when you want to go deep on a topic, even if the router thinks it's an easy question.
skybrian•Aug 6, 2026
Some people want quick results. Some people want it to keep searching for a new math proof overnight without giving up, and they have money to burn.
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
pllbnk•Aug 6, 2026
The _Intelligence_ part of AGI should be able to guide the user through that without all the knobs.
vanuatu•Aug 6, 2026
intelligence is not omniscience though
stymaar•Aug 6, 2026
Asking questions to clarify user intent is a very low bar for intelligence. A bar that all SOTA models fail consistently at though. (It's both funny and legit infuriating when Opus, after having made a dozen wild assumptions without checking with you, then comes back with a request for clarification on some mundane topic).
vanuatu•Aug 7, 2026
but you can keep asking questions ad infinitum
very nontrivial problem knowing when to stop and making assumptions
minimaxir•Aug 6, 2026
OpenAI tried auto-routing with the initial GPT-5 release and it was immediately clear why that was a bad idea.
dbbk•Aug 6, 2026
I don't understand this at all. Whenever I ask Gemini 3.1 Pro Extended, or Claude 5 Max something in chat, the most I ever wait is maybe 30 seconds. Is that really so bad?
oceanplexian•Aug 6, 2026
"Wait 30 seconds" as a concept has been totally incompatible with the web, smartphones, etc for about 20 years now.
awakeasleep•Aug 6, 2026
Because your incentives are opposed to the provider’s incentives
redox99•Aug 6, 2026
Because the model can't read your mind and know if you want a quick answer, or an hour long deep dive.
Jtarii•Aug 6, 2026
Then it should just ask the user what they want if its unclear from the context.
2sk21•Aug 6, 2026
Exactly! I posted much the same comment in another thread and there were lots of huffy complaints that amounted to "you're prompting it wrong"
Sammi•Aug 6, 2026
That's what the reasoning slider is for!
redox99•Aug 6, 2026
So instead of just getting an answer, I have to wait until it asks me, and I have to type back a response? Extremely annoying.
egorfine•Aug 6, 2026
I absolutely need instant mode as this is what I use 90% of the time.
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
nojs•Aug 6, 2026
Auto-effort and similarly auto model routing suffer from a halting problem sort of issue: you don’t reliably know if a request is complex unless you use a complex model to make the decision.
miki123211•Aug 6, 2026
I have the opposite problem. I'm not well-calibrated on when I'd want lower reasoning than what's available to me (and how to compare that to lower-tier models). OpenAI now has Luna, Terra and Sol, each at Low, Medium, High and Xhigh, with Pro/Ultra depending on harness and plan. That's ~15 possible combinations of model and reasoning level, and there isn't a satisfactory explanation of which one you want for any particular task.
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
redox99•Aug 7, 2026
Yeah, it's hard, there's so many permutations. But most are bad so it narrows it down.
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
drivebyhooting•Aug 7, 2026
Why not pro/ultra?
redox99•Aug 7, 2026
Pro: It's web chat only, I haven't tried it so idk. I think it's useful for Math and that kind of stuff, but I only do programming.
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
drivebyhooting•Aug 7, 2026
It’s very confusing to me.
I’ve occasionally used pro to do high-level research and design. Then I ask it to create a prompt for Codex ultra. Ultra can do a lot of genetic benchmarking and testing to elucidate and resolve quandaries.
I don’t know if this is a good workflow.
majormajor•Aug 7, 2026
> (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High)
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
customguy•Aug 7, 2026
just wait for the loot boxes, mark my words
kingstnap•Aug 6, 2026
Its always fun to try to read between the lines here to speculate why they are doing this.
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
deanc•Aug 6, 2026
It's also going to be more training data for them.
ToValueFunfetti•Aug 6, 2026
Is this new behavior from them? I haven't looked at their free chat offering in a minute, but I thought they always had them close behind paid tier, often with essentially equal products that made it weird for them to sell the paid tier for chat.
mkozlows•Aug 6, 2026
No, free tier was total garbage with GPT-5.
kingstnap•Aug 6, 2026
You can't even pick 5.6 luna on the chat app with a paid subscription. It just gives you sol (or use older models) with what seems like basically as much usage as you want. And sol is considerably smarter than luna.
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
simianwords•Aug 6, 2026
> Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
planb•Aug 6, 2026
Let’s speculate: they want luna to be their first model “on silicon” and need as much test data as possible before finalizing the design.
porridgeraisin•Aug 6, 2026
I dont pay for a chatgpt subscription, but sometimes I did use the web app for throwaway questions. GPT 5.5 Instant or whatever it was that they had was absolutely horrendous. Never answered a question straight and was pedantic in a way even a redditor wouldn't be. So I dropped it and just opened my paid coding agent for everything. Grok.com is quite good now with grok 4.5 though and I find myself using that often. Hopefully luna will be similar.
ilaksh•Aug 6, 2026
> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users.
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
kkoncevicius•Aug 6, 2026
If we could show the current models to someone like Alan Turing, I am sure he would conclude that we have AGI.
stymaar•Aug 6, 2026
And after half an hour using it he'd just admit that his test was way too simple as these models are still way too dumb
klibertp•Aug 6, 2026
I think of the Turing test as one of the starting lines, along with image recognition ("a summer break project for a group of grad students" resisted being solved for decades).
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
stavros•Aug 7, 2026
How are the models too dumb? How are they dumber than the average person? I wonder if anyone gave Claude an IQ test (the one for humans).
ilaksh•Aug 7, 2026
Yes there are a few sites with IQ test benchmarks. The frontier models come up around 130 or 140 depending on which model/test.
stavros•Aug 7, 2026
That tracks, I wonder what the people who say that the models are too dumb expect to see. Miracles?
hluska•Aug 7, 2026
Dumb is too simplistic, but they sure don’t pass for being human. That would be the intent of a Turing test… not IQ.
kubb•Aug 6, 2026
> This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Well, the models are smart enough to point out why this is wrong.
lostmsu•Aug 7, 2026
They were trained to do that specifically, don't you think?
andai•Aug 6, 2026
>It also needs to be differentiated from ASI with godlike powers many times greater than human.
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
ilaksh•Aug 6, 2026
Which brings up another point in that we should have a term that distinguished between super intelligence in some capacity and godlike super intelligence.
taytus•Aug 6, 2026
"This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud."
I cannot believe the comments I read on HN nowadays. ZERO critical thinking.
lukevp•Aug 7, 2026
Well your comment adds no value and is just offensive. Why don’t you argue a counterpoint?
aniceperson•Aug 6, 2026
a m-dash in the title? whoa things are degrading fast
heaney-555•Aug 6, 2026
Giving free ChatGPT users access to reasoning (the 'Think' toggle) will have a broader impact on the world than every new paid model and coding agent combined.
daemonologist•Aug 6, 2026
Going from 5.5 Instant (which was noticeably bad) to 5.6 Luna is a big jump as well. OpenAI is probably the most prominent among the general public - an advantage in some respects but they're giving away a lot of free inference and thus have to use a pretty small model to do it.
johnsmith1840•Aug 6, 2026
I really don't see how they're going to be google here. The free tier is dominated by verticle integration and the platform that people use. Long term I imagine google wins the bottom of the market and I'd be suprised if they lost.
Tostino•Aug 6, 2026
People tell these models everything. You don't think a pretty unethical company can figure out a way to monetize that?
fragmede•Aug 7, 2026
If someone tells ChatGPT I'm cheating on my wife, sure, unethically they could blackmail the user but that seems like a stretch.
hluska•Aug 7, 2026
I think the consumer profiling side is more compelling (and likely) than blackmail.
Tostino•Aug 7, 2026
Yup, that is where my mind was.
tokioyoyo•Aug 6, 2026
Ads.
stymaar•Aug 6, 2026
Ads work when you can server $.01 worth of ads to a user for $.0001 of server cost. I fail to see how you can make it work when doing LLM inference which is significantly costlier than web search.
tokioyoyo•Aug 6, 2026
In about 4 months, OpenAI’s fundraiser decks will leak, which will include their current ad revenue.
livinglist•Aug 7, 2026
Is this a promise or speculation?
kristofferR•Aug 6, 2026
90+% of queries are probably so common/evergreen that, if a cheaper model made all the different language varients and ways of asking the question into one query, they could be cached quite effectively.
virgildotcodes•Aug 6, 2026
I recently saw some discussion about the CPC for personal injury lawyers being ~$200 on Google ads. Seems insane to me, but I’m totally disconnected from the advertising world. That said, depending on the conversion rates between chats and clicks, the math could be favorable for OAI.
mikeshi42•Aug 7, 2026
High CACs are supported by high LTVs, that is about the right ballpark for things like lawyers or even plumbers.
famouswaffles•Aug 6, 2026
> I fail to see how you can make it work when doing LLM inference which is significantly costlier than web search.
The median LLM query isn't significantly costlier than web search.
fragmede•Aug 7, 2026
Except that Google is giving AI answers for ~every search now, so they're paying that cost for every search now as well.
keeda•Aug 7, 2026
That's a problem Google has to solve too ;-)
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
johnsmith1840•Aug 7, 2026
More like LLMs are near commodity at the common tier the ones who wins are those who can inference the cheapest and google is easily thr best positioned to do that with many years of custom chip model optimization.
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
ramraj07•Aug 7, 2026
Even a long while ago it was estimated that every Google search cost google several cents (and that they made back 10 cents). Google also uses AI for search results. Thus its not inconceivable that openai can make money here.
johnsmith1840•Aug 7, 2026
I typed wrong I mean I don't see how they will beat google. Google is the only true vertically integrated AI company meaning they own price at the bottom margins with their own chips and models made just for them.
ekidd•Aug 7, 2026
Google's public AI model lineup is pretty bad right now. Gemini 3.1 Pro is scoring worse than some mid-sized Chinese models at 1/20th the cost, and their Flash and Flash Light models are horrendously overpriced on task benchmarks compared to GPT 5.6 Luna or the mid-sized Chinese models.
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
gavinray•Aug 6, 2026
I'm not sure people fully comprehend the trickle-down pop-culture/zeitgeist effects that LLM's are having/are going to have on humanity.
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
crab_galaxy•Aug 7, 2026
Is that really accurate though? I use Claude in my software job but the majority of my friends do not use it at all. And I never use it for anything other than that either.
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
fragmede•Aug 7, 2026
We're all in our own personal bubbles and extrapolating from there as normal, but ChatGPT has some 800 million users, most of them outside the US. Countries like the Philippines and Thailand. I don't think they're all vibecoding the hottest new SaaS app but maybe they are.
thorum•Aug 6, 2026
Free users have had access to reasoning for a while. o4-mini and the initial GPT-5 launch both included reasoning modes for free users.
They took away the button a few months ago and are now putting it back.
heaney-555•Aug 6, 2026
Technically yes, but the free usage equated to single-digit messages per day before the auto-downgrade. With GPT-5 it was just 1 message per day.
Perhaps I should have said "proper access".
in-silico•Aug 6, 2026
Does the free tier really not have access to reasoning models?
That would explain a lot of the terrible AI/LLM takes online.
heaney-555•Aug 6, 2026
The free tier of ChatGPT is powered by GPT-5.5 Instant, a non-reasoning version of GPT 5.5.
fragmede•Aug 7, 2026
Yes. Go to chat.com in an incognito window and ask it the same question that you ask 5.6-sol with thinking max, and compare. Depending on the question there is or is not a huge difference.
qingcharles•Aug 7, 2026
Also, remember the bulk of AI users are using free models, getting terrible answers, and wondering why people keep saying AI is going to take everybody's jobs away.
simianwords•Aug 6, 2026
Does 5.6 Sol finally have an "instant" form? Is that the change?
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
simonw•Aug 6, 2026
It's SO hard to understand this. I couldn't confidently explain it at all.
hluska•Aug 7, 2026
It’s not subtle when you point it out - you’re looking for a different phrase entirely.
OsamaJaber•Aug 6, 2026
Expanding free access is mostly an inference cost
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
Squarex•Aug 6, 2026
Why pay for ChatGPT Go then?
HDBaseT•Aug 7, 2026
Access to "ChatGPT-Image-2" I suppose?
sunaookami•Aug 6, 2026
This is actually a downgrade for free users since currently it uses GPT-5.5 for a few messages before it drops you down to GPT-5.5-mini. Now it always uses a model worse than Mini (Luna is nano-equivalent, "It roughly corresponds to the nano model tier used in earlier GPT-5 families." https://developers.openai.com/api/docs/models/gpt-5.6-luna ). I guess it's a bit better with Thinking though. They should use Terra for a few messages first before dropping down to Luna. And image inputs are still limited.
saithound•Aug 6, 2026
While their math results are impressive, vibe coding their own web UIs with their subpar design models is really going to backfire if their plan is to attract new users with better free model offerings.
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
applfanboysbgon•Aug 6, 2026
> Our mission is to ensure that artificial general intelligence benefits all of humanity. We’re introducing updates to ChatGPT that improve everyday conversations while expanding access for Free users.
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
kgeist•Aug 6, 2026
>avoid extra detail when it does not help
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
kaszanka•Aug 7, 2026
> Because this version of GPT‑5.6 Sol is optimized for everyday chats, it will only be available in the Chat experience in ChatGPT. The version of GPT‑5.6 Sol that powers Work and Codex is not changing as part of this release.
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
johnnyApplePRNG•Aug 7, 2026
Excellent.
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
roytam87•Aug 7, 2026
OpenAI should be more transparent to (free) users about Data Analysis quota and Chat with attachment quota.
1saadcodes•Aug 7, 2026
I dont get why they'd make this available to free users when they're probably dealing with compute constraints considering their competition with the Chinese models. Giving a more capable model to a huge free user base looks like an expensive choice to me
freakynit•Aug 7, 2026
They've probbaly have realized that the foundation models are becoming commoditized (or will be very soon). We can already run luna medium class models on phones now (bonsai, gemma4).
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.
23 Comments
luna is very good
Both models are cheap enough that I can run 4 sessions at the same time without running out of the 20 USD codex and 10 USD Opencode plan. I've burned through almost a billion tokens this week and I've done some pretty big refactors as well.
I have a Claude Max subscription but I've barely been using it because of the many issues they've had this week.
Pricing absolutely matters.
But seeing the graphic with the visual weather report: that makes me think that is not the goal at all. :)
Even after identifying the 0% chance of rain, it still drags the conversation on and on and on
I expect a few things to happen in the next year:
1) Exclusive MCP server deals/API integrations
2) Significant switch to B2B marketing, even moreso than we've seen before, with API interfaces being paid and chat-client interfaces becoming more and more free, perhaps just with limits more on integrations or data visualization/analysis
3) US restrictions on B2B contracts with non-US hosted models that do any sort of contracting with the government
Obviously there's a bunch of stuff I'm not foreseeing. But it really does feel like the bottom of the market is collapsing into free. I assume OpenAI and Anthropic think their next generation of models will restore their halo tier status and that the cash burn is justified to just get there, but this has to really mess up IPO plans.
1) Back then, even as a free user you'd be able to use the strongest model (even if with tight limits). Now, you need to pay to use Sol, and you need to pay to use Opus or Fable. It does seem fairly premium in that sense. Idk about 5.6 Luna, but the previous Instant was really bad, even for very casual users. It would hallucinate non stop.
2) When $100 and $200 per month plans launched, they were received as outrageous even here. Nowadays they are pretty common among power users.
This is kind of what I'm saying though. Bottom has fallen out, differentiation is just can you be much more premium than the competition. Currently that remains unanswered.
EDIT: I'm basing this off the assumption that for chat, premium is not a point of differentiation at all. For coding/analysis, it is.
Back when coding for me still meant copy-paste from the web version, it was only worth the $20/month for me.
They only added the $100 Pro plan in April during GPT 5.4 times.
Today I happily pay $400/month for Codex and Claude Code.
[0] Won't put it past OpenAI (and/or Google) to open weight larger models!
So 5.6 Luna is just their next version of what they used to call 5.5 instant tier
And 5.x instant models were never much to write home about anyway so the default ChatGPT free model hasn’t been particularly distinctive since 4o
It seems like giving it a time limit or a budget in dollars would be clearer, though?
Or, keep searching until I come back to the computer and ask about progress.
very nontrivial problem knowing when to stop and making assumptions
Sadly it's not available anymore on the updated desktop app (formerly Codex) and the previous desktop app (formerly ChatGPT) is abandoned.
I feel that work is basically split into two tiers, hard (which requires a good model and lots of reasoning by definition) and relatively easy (which won't consume much of my limits despite a great model and reasoning, so I may just as well keep it on Sol High).
Luna: Can be useful if price sensitive, always use with AT LEAST high effort. But codex plans are very generous, so just ignore it.
Terra: Forget it exists
Sol: Just use this. Medium can work for specific edits. Otherwise just use high or xhigh.
OpenAI basically agrees with this, and the slider gives you those options.
TL;DR: Just use sol medium/high/xhigh
Ultra: Might be useful, the one time I tried it, it was extremely wasteful in terms of tokens. Definitely not a daily driver.
There's also max effort, I think it can be useful but the gains compared to xhigh are quite small.
I think max and ultra are only worth it when the others fail.
I don’t know if this is a good workflow.
This equation significantly changes if you're paying API prices vs on a subscription. (Such as if you're integrating it into a different product, where it becomes worth it to figure out what the cheapest you can leverage is.)
Maybe Luna efficiency gain was actually significant enough that putting all the free users and giving them super generous limits makes sense.
They might be doing this to improve the messaging of AI among causal users since right now there is a huge amount of datacenter backlash in the US due to AI grievances.
Maybe they have too much excess capacity or they really want to juice token numbers and market share on their dashboards for marketing.
I also wonder if being given access to an actually a decent model like luna with actual thinking budget instead of brainless "instant" modes will start to make causal users understand the real capabilities of these models.
All of this is of pretty minor importance though. You can't read as many tokens as a subcription can produce so more chat is not the value add nor super important.
I mean there are literally so many providers for free chat if you are willing to use several seperate apps.
The real value in these subs is using codex cli, much like the real point of anthropic subs is using claude code. Because agentic work actually does require a lot of tokens.
Definitely this. The recent 80% discount was a reaction to Deepseek's update so that they still position near the frontier. My theory: Luna has always had a much higher efficiency. You do know that the model didn't get faster after the discount?
This clearly implies that they believe ChatGPT models are AGI and are now willing to say it out loud.
Which I think is a fair interpretation of the term. They are general purpose intelligence in that you can get help from them about almost anything. They are not like narrow single purpose AI models.
I don't think we need that term to mean "can completely emulate a human" or "can do every task any human on earth can do as well as them".
It also needs to be differentiated from ASI with godlike powers many times greater than human.
It is a huge leap. Now we can start talking about "intelligence" at all - we really couldn't before. That we're still hovering barely above the starting line is a separate matter (also worth noting, of course).
Well, the models are smart enough to point out why this is wrong.
Wasn't there a thread the other day about how very few humans on earth can understand the new math proofs?
Although "with sufficient study" vs "not even with unlimited study" are probably worth distinguishing there.
I cannot believe the comments I read on HN nowadays. ZERO critical thinking.
The median LLM query isn't significantly costlier than web search.
To expand: it seems inevitable that Google's SERP format will be replaced with a conversational / chatbot / agentic interface, which equalizes the playing field for all chatbot providers.
This is because you can stuff in much fewer ads into a chat interface compared to SERPs. (They could try stuffing more ads but that would likely just push users more to the competition who have a much lower baseline on which to show growth.) As such, Google would be forced to progressively nullify its own invincible firehose of ad revenue as they deprecate SERPs in favor of AI overviews.
Google is going to be able to push cost down further and for longer than anyone else. They also rightfully believe this is a company ending gamble if they fail they could be a second tier player for decades.
So better inference margins, far better operation experience serving cheap AI at massive scale, a massive warchest, and the fear of complete company irrelevence.
That's going to be one hell of a company to beat for free tier LLMs. Also chatgpt is impressive but it's notthing like google search quite yet. On top of that AI labs must use google's product for their AI.
They also own more data by an exponential margin. Anything but dominance of free AI on google's side would mean they are just so incompetent they deserve to fail.
There's no reason why Google's public stuff is this stale, overpriced and underwhelming. But at least until their next round of models drops, even calling them a "frontier lab" is starting to feel like a stretch. Which is weird!
Because everyone now outsources much of their thinking and researching to LLM's, our collective culture + brain is shaped in a cyclical manner by using them.
It's the mechanical homogenization of culture and groupthink.
TBH I don’t find it useful at all for personal use. It’s totally soulless for creative ventures and absolute dogshit at researching the things I want it to be good at (I.e. planning a vacation or finding new music).
All that combined with the social stigma makes me feel pretty skeptical that it’s some kind of pop culture shaping mechanism, at least not for a few more years.
They took away the button a few months ago and are now putting it back.
Perhaps I should have said "proper access".
That would explain a lot of the terrible AI/LLM takes online.
(This is a subtle nudge at anyone from OpenAI who reads this to make sure they get updated.)
OpenAI have a model called "chat-latest" - I wonder if that's running this new model yet: https://developers.openai.com/api/docs/models/chat-latest
It's described as "points to the latest Instant model currently used in ChatGPT" - so presumably that's "GPT-5.6 Instant" in the app.
https://gist.github.com/simonw/aae4febd3c6f7bc5b7811857edb3c... has screenshots that still show "Instant" as an option for ChatGPT Chat... but not for ChatGPT Work.
It’s hard to understand this .. like sure we can select instant but is there an actual model called 5.6 instant? Like is it on LM Arena and OpenRouter or available via API etc
5.5 instant is definitely A Thing it’s even name checked in this OAI post
Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
The Aug 6 update has forced the entry box to auto-format Markdown in an attempt to imitate Claude. The implementation is buggy and even simple copy-and-paste has gone entirely haywire. They also forgot to leave a switch to turn the confounded autoformatting thing off.
Chat mode in general is currently crawling with more UX bugs than a porch screen in summer.
My mission is world conquest. I'm writing a comment on an HN thread.
No, those two clauses have no relation whatsoever. I just felt like saying the first sentence because it sounded cool.
I wonder if they actually do it to optimize inference. I maintain a corporate AI server and one of the tricks to reduce the load was to modify the system prompt to be as terse as possible so the average response completes faster and requests queue up less often.
Hm, does this mean that 5.6 Pro in ChatGPT web is somehow different/not as good now? I found it really good for code review (upload your repo and patch and off it goes).
Yes, please give free users more access and leave paying codex users in the dust.
Wise plan, Sam.
The value is shifting up the stack... make core intelligence free, then monetize the ecosystem built on top of it .. kinda similar to how the internet itself is free, but platforms and apps capture the value.
I think we'll see a huge push towards connectors for work and personal tools, along with much deeper os level integration. That's where the long-term moat is, not the base model itself.