GPT-5.6 --------> GPT-5.6 Sol
GPT-5.6-mini ---> GPT-5.6 Terra
GPT-5.6-nano ---> GPT-5.6 Luna
Two important things to note, if you want to verify what I say/correct me:
GPT-5.6 Terra actually scores worse than GPT-5.5 on many benchmarks. It's not GPT-5.5 trained with more compute; it's basically GPT-5.6-mini that's been distilled from GPT-5.6 full size. Remember, GPT-5.4-mini had almost the same benchmarks as GPT-5.2 after all.
Opus 4.8 runs at ~90 tokens per second. Fable 5 runs at ~40 tokens per second on from Anthropic, because it's a bigger/slower model. A few days after the release, when the dust dies down, look at how many tokens/second GPT-5.6 Sol is running at. I will bet it's the about same as GPT-5.5, and not half the speed. (OpenAI is not incentivized to slow down the model for paying customers). But the model tokens/sec will be a big clue- if OpenAI is charging more money for the same sized model or not.
And the "business" obvious is still doing that but the science and implementation has be realizing that this just isn't true. They're not getting AGI out of a single LLM by itself.
There's a lot to make efficient, but it should be clear to everyone that just throwing compute at larger models isn't going to magically make it rain.
What is this very confident assumption based on?
> It's a damn good model. Not quite as "smart" as Fable, but it is incredibly capable. Fixed all the problems I had with GPT-5.5.
> It is incredibly determined. Will run for a day without even using a /goal. It understands subagents incredibly well and is great at orchestrating. It's super pleasant in use cases like OpenClaw and Hermes Agent. It knows iOS dev incredibly well.
> It has rough edges too, but FAR fewer than 5.5 did.
> For many things, gpt-5.6-sol will become my obvious defaults.
> It is better about [following instructions] than 5.5 was. Understands intent well and hammers until it gets there. Sometimes a bit too hard.
Also[^1]:
> gpt-5.6-sol is world leading in computer use. It made me use it 100x more. When we lost access to 5.6, I quickly started to go insane without it
[^0]: https://nitter.net/theo/status/2074708892341481755 [^1]: https://nitter.net/theo/status/2074720467395756499
Every time I've ever seen one of his videos it's pretty clear he has very little understanding of development or engineering. I first became aware of him from his early "unit tests are a waste of time" stuff, and it seems his skillset is building a personal brand. Fair play, he's clearly talented at that, but that doesn't make his opinion on anything else worthwhile.
I cannot prove it but I have a feeling that you may be conflating "he clearly has different opinions on things I consider non-negotiable" to "he doesn't know what he's talking about".
I also watched a lot of his videos. I wildly disagree with him a lot of times, but he has his reasoning, and I can see (and verify!) that those ideas are coming from an engineering perspective.
That's his motivation, influencing. Not teaching.
I'm not so against him as the previous commentor but I feel basically all YouTubers who have succeeded in building a brand have the same problem. They have to present their opinions as unassailable truth, they can't allow nuance. Within reason of course, they also have to play the game of appearing considerate and understanding of other people but it will always boil down to proving they are the real experts, their ultimate goal is always to get you watching more of their content.
If they don't do this, they appear less trustworthy and their brand wouldn't have grown as much as it did. They might genuinely have some expertise to share, and even contrary or downright wrong takes could teach us something if they were only presented in a way that encourages critical thinking. But as the person above said, their real deep expertise is in brand building.
Of course, youtube isn't interactive and when you see something that you think is objectively wrong, your options are writing a comment nobody will read or ignoring it, which is frustrating, but that, in my opinion, doesn't discredit the content producer itself.
Maybe you are desensitized to it, but I have a carefully curated YouTube and I know that it can be a completely different platform experience if you reject with prejudice any such content.
And this goes back to my point, if you are a competent engineer, why are you spending time producing rageslop for money, rather than, you know, doing the engineering.
I don't agree that he has demolished his credibility. I also dislike the youtube face and sensationalization but I personally don't hold it against him, given the Youtube algorithm.
Regardless of his style, I like hearing the take from an engineer who's working in a different country/culture and has a completely different perspective.
edit: it seems you changed "demolished" to "harmed". I still don't agree but it reads more defensible IMHO, thank you.
His youtube channel used to be about talking about the new FOTM Javascript framework/technology - not presented as 'here's a cool thing, let's check it out' but 'everyone worth a damn already uses this, get with the times grandpa'
It's shocking how many accurate tropes this hits.
I don't get many programmer influencers in my feed that deal with newsworthy relevant stuff. Theo is the least wrong and most humble one in my perception.
None of these were Theo's take. He was pushing the idea that unit tests in general were a waste of time because you could be shipping new features instead.
https://www.youtube.com/watch?v=pvBHyip4peo for an example of this. The nicest possible interpretation on this is that he's deliberately saying something he knows is wrong to attract attention.
1: get bug
2: write tests that should work, but don’t because of bug
3: fix bug
4: confirm fix by running tests
Makes things a LOT easier for people checking the PR, they can just confirm the tests are correct pretty much.
As a bonus the same bug can’t surface again.
I think the value is much lower (maybe even negative) when you're still trying to work out what shape the code will take, in an initial implementation.
Of course, as others have pointed out, nuanced opinion doesn't get clicks or YouTube views.
But you're right, the goal is not to write test but to ensure delivery of a reliable software. However each software is a prototype, something that has never been made before (unlike a manufactured car or chainsaw) so the customer must be ready to some unexpected behaviors when the software is released.
Since tests are often sloppy or does not cover every edge case, I see a real value for GenAI. It also forces to write good spec: very clear about inputs and the invariants for each use case. I think that AI (especially GenAI) should first be a solution to existing problem, lack of tests and good specs is often one of them.
(A little toning down of the goblin fetish would be nice too, haha.)
I looked at his YouTube, and found a stream of industry gossip and beginner content like "web dev tutorials". I have nothing against such content and it may be useful and good fun to watch.
But does that say anything about this particular model? People have been using models effectively for web code since Gpt 3.x.
THIS IS BECAUSE GPT-5.6 SOL IS... just a more posttrained version of GPT-5.5, not a brand new bigger model than GPT-5.5. It's not like how Mythos is bigger than Opus.
OpenAI switching to Sol/Terra/Luna renaming is just a way to rip off people and charge more usage for the same sized model.
GPT-5.6 --------> GPT-5.6 Sol
GPT-5.6-mini ---> GPT-5.6 Terra
GPT-5.6-nano ---> GPT-5.6 Luna
Except OpenAI is about to advertise GPT-5.6 Sol and GPT-5.6 Terra as a whole tier better, than if they named it GPT-5.6 and GPT-5.6-mini.
So, if they improved a lot in those areas, then GPT-5.6 could become a lot more useful compared to GPT-5.5 even though it might score lower in many benchmarks. It's possible but unlikely since their approach was mostly brute force in the past.
Claudes are more creative and get shit done, suggesting and implementing stuff you didn’t ask for but actually kinda needed. Will leave gaps and bugs though. More of an artist, communicates a bunch during the dev process too.
GPT is the engineer, given exact specs it’ll disappear into its dark corner and putter away at doing exactly what was asked, nothing more nothing less. Very very good at spotting gaps from Claude’s get shit done code.
Fable 5 meanwhile has a reliable 1m context window and compaction that the few times I did eval it does also do well. Not quite as easy to trust as GPT-5.4, but that's mainly because with thats 272k context window I simply got more familiar with GPT-5.4s incredibly dependable compaction.
Purely concerning encoded information wise, Fable 5 is near or on the same level as Gemini 3.1 Pro in my limited test set focused on those tasks, which in very niche cases can make a difference even with coding, but the truest advantage for coding assistance (besides frontend/UX) is that the code Anthropic models provide is more parsable. Hard to explain, but I can read, follow and mentally map Fable 5 (and even Opus 4.5-4.8) output far more than GPT-5.4 or GPT-5.5 code.
Task orchestration and (more importantly) knowing when to recommend against using such vs Opus 4.8 is another strength of Fable 5 I've use liberally, there is an understanding of what a tasks requirements and the most optimal setup for success are, I have not yet seen before. Computer use is also solid, albeit not as token efficient as GPT-5.5 for my limited use cases.
Lastly, I will say that the classifier has become far less intrusive for me compared to the initial release. During the previous launch window, on Claude.ai I triggered the classifier for simple frontend tasks for regular (not security vocabulary containing) webpages. Now that is no longer the case. Inside Claude Code I occasionally triggered the classifier previously, but after the re-release, I only managed one, even when working with a privacy focused section of the code base containing a significant number of code comments with security and privacy focused wording. That one instance was rectified quickly by trying again, so I really am having a hard time following how others experience the issues some describe. I do have routing to Opus 4.8 without confirmation by me deactivated too, simply because I want to know if it ever happens, so it's not that I missed reroutings.
That all being said, we are still far from a stage where I'd not want to review the output, but yes, I do rate Fable 5 very highly. GPT-5.5 can have a similar ceiling but long horizon has become less usable over GPT-5.4 and in either case, parsing their output is (far more) of a chore. Maybe post training can address some of this, hopeful on the compaction front myself. Also interested in what happened to OpenAI models on AWS Trainium, I was expecting that to be a major boon for their commercial adoption, but haven't heard anything since then...
On the post training front, I am still hopeful that the Gemini team can finally get tool calling and task adherence to an acceptable level as we do need every competitor possible and purely considering the information density the model was trained with, they have great potential.
Excuse me, but what are you on about?
Unless I'm mistaken, they have literally(1) stated that it will cost $5 per 1M tokens in, and $30 for 1M output tokens. The same as GPT-5.5.
Mythos is simply a much bigger model in terms of parameters and I don't think OpenAI will have anything of its size anytime soon (My theory is that OpenAI had given up on scaling up parameters after GPT4.5 flopped).
> We generally treat GPT-5.5’s safety results as strong proxies for GPT-5.5 Pro, which is the same underlying model using a setting that makes use of parallel test time compute.
And Gemini also provides something similar. Gemini Deep Think models are pretty much the same thing [2]. As to why no other company uses this, I don't really know. Maybe compute constraints?
Spanish: Sol, tierra, luna
Italian: Sole, terra, luna
Catalan: Sol, terra, lluna
Portuguese: Sol, terra, lua
Might as well call it gelatto, siesta, fiesta if they think it sounds cool.
https://x.com/theo/status/2074708892341481755
5.6 sol seems to hit a lot of the gaps with 5.5
sucks its not "mythos" but i will take it
I will stop here, sorry but I think we have limited time to listen to opinions and nowadays since they are abundant on social media we should give preference to the substantiated ones.
If there’s anything I learned over the past 12-18 months is that this is a recipe for disaster, except for throwaway stuff.
I thought most senior engineers settled on the fact that steering a model yields much better results?
Just last night Fable decided to get into a rabbit hole of debugging a database driver issue by packet sniffing the network traffic instead of just adding debug statements to the code. Definitely needed steering, and I don’t know many people whose first intuition would be to use pcap when they have a segfault.
In my experience even Fable still requires guidance (although the options it provides are generally better).
I also flip between the models due to quota, TUI enhancements, model updates and service availability.
To handle this, I built a thing that normalizes your transcripts between Claude Code and Codex into a shared DB, then a CLI and skill.
It has made it so it doesn't matter what I built where (or when) I just refer to the work and drop in a /total-recall (or $total-recall on codex) and the agent brings it into the current convo.
I realize there are a lot of ~memory tools out there, but I think particular my approach and product behavior is unique.
If you're open to giving it a try, I'd appreciate any feedback: https://contextify.sh recent show hn: https://news.ycombinator.com/item?id=48777790
You are the sole owner of the project implementation.
User maintained documentation:
- goals.md for the project overall goals
- tech.md for guidance on how to build the project
Agent maintained documentation, current state living specs, these are not logs:
- project.md is a map of the code, components and features.
- choices.md write here all decision taken by the user.
Do not duplicate information between these document.
Personally I just , in the orchestration loop, have all decisions be constantly reviewed and deliberated on and the decisions logged in a permanent way, that way everything is auditable, the model if needed can go back and look at why x or why decision was made or x or y tool used, and they're all labeled as D-1234 or whatever.
Plus I have it log the council discussions and always include provenance or the opinions so fable can go back after every major implementation and review how the orchestration loop could be improved. Basically have it log as much thinking in an organized compartmentalized way is better than any memory feature I've found though I haven't tried many. Auditable logs for every major decision, use 5.5 with reckless abandon (still have 3 resets myself).
Not claiming this is perfect but it has led to a very easy time of any fresh agent picking up the project. I also keep a task queue and project status and agent playbook that also get refined based off the logs of how a run went
e.g., it still doesn't have /revise or /undo!
My agent has access to glab with a user and can do whatever within permissions. No need a MCP. MCP maybe just for browser control.
I canceled Claude plan a few months back and have been using this. OpenAI plans are much more generous.
Edit: I guess the real answer is that I don't want multiple subscriptions. I use one for a month, and then decide if I switch to the other.
What type of work is this for in your experience?
How does this go through so many layers of management at a trillion dollar company without who has a say raising this? I simply can't believe how stupid the naming scheme from OpenAI was and continues to be even after they acknowledged it earlier.
Possibilities are endless:
- OpenAI GPT 5.6 F5+
- OpenAI GPT5-6000FT
- OpenAI GPT5-6505FS
- OpenAI GPT5-6F05UL
Sounds nice, looks cool. Why not?Past Intel had way better naming.
It'd just be OpenAI GPT 5.6++++++++++++++
OpenAI GptForce 5600 XT
OpenAI GptForce FX6800TI Founders Edition, for example.
On a more serious note, I can vividly imagine how difficult it is to agree on a set of words that could plausibly suggest a relational meaning while remaining non-diminutive in every individual model name. Adding to the complexity, it is going to be used globally, and the main competitor already has an arguably successful, fabulous naming scheme.
It sounds like a PR minefield.
What surprises me is not this, but that OpenAI changed things up without syncing with a GPT 6.
It sounded nicer than something like "Luxury" and "Basic".
That’s interesting. My understanding is that:
Sol = Sun Terra = Earth Luna = Moon
So it’s a bit surprising that in Toyota’s nomenclature, Terra is the basic trim instead of Luna.
(Well given the limited amount of things we can deduce from a name)
mini > micro (see, e.g., -skirts, -computers)
micro > nano (see, SI)
so, mini > nano.
Sounds like the source of your problem, right there.
Whereas Sol/Luna/Terra reads more like "GPT for hard/medium/basic problems".
Another dimension for the fronteir to move in is speed. Codex has /fast which is great, but yea the bottleneck right now in many cases is just the time it takes these tools to complete tasks. I'm running many sessions in parallel just because I'm waiting for tasks to finish. I'm constantly round robin'ing them, and kicking them off on the next 20 minute task. If these models were faster I wouldn't need to context switch as much.
I sometimes chuck a few tokens to gpt 5.5 and opus 4.8 and they can sometimes solve a problem one of the other models couldn’t, but they’re not like 10x better or anything in my experience. More like 1.2x better
If all I had was Fable for the next couple of years then I'd be totally fine with that. I have never felt that about any version of Opus.
Using cheaper models and using your skills and expertise from the pre-AI era can get you working just as fast. You've gotta be more specific about the work you need doing. It's less "vibes" based, but they're still effective.
Also, Chinese models absolutely are taking off. I used Claude and GPT at work, and then I tried using some Chinese models for personal projects. I am 100% convinced they're like 90% as good for 10% of the cost. But you've basically gotta be a good developer first and know what you want and know when it's giving you shit.
Of course, if you think that this approach is as fast and effective as "vibe coding" as in outsourcing more thinking to the AI, it is not surprising you would conclude the cheaper models were nearly as useful.
I don't know if you are right or not, a lot depends on the constraints of the project and team.
Some of the newer arguably now viable use cases, such as porting a large codebase to Rust, are certainly not going to be as fast with a more manual approach.
I’d describe it as something between Sonnet and Opus.
Eventually it plateaued and now you can do a decent chunk of your computing on something from 2012.
People keep saying scaling will top out, for example. But scaling keeps stubbornly refusing. New techniques keep coming along too. It's really still exploding into existence and every new generation brings new capability. Eventually it'll clear a ceiling for your key use cases and you'll stop worrying about new models.
It always pays to look back at history and see if you can pattern match.
If OpenAI can launch a Fable tier model that's actually usable on a subscription, then Anthropic is just going to lose, and badly.
Same also for the announced changes around `claude -p` and Agent SDK use that were backtracked
Codex CLI just seems faster at coding than Claude Code but Fable is just a level above intelligence wise, it's truly like taking to very very very smart human.
With GPT 5.6 though will be interesting to see if things flip, to have Codex speed (or faster) with Fable level intelligence is a game changer.
Though its been just 3 days I started using.
Half way through the chat, GPT 5.6 Sol stops and does a safety verification, pretty annoying
I'm guessing this works better because it can always go back and re-analyze the saved context.
Our docs show a diagram here:
https://developers.openai.com/api/docs/guides/reasoning
> Input and output tokens from each step are carried over, while reasoning tokens are discarded.
Keeping reasoning tokens around is better for caching and for remembering past insights, so you might reasonably wonder why we designed it this way. The main benefit of dropping reasoning tokens is that you can fit a lot more work inside the model's context window before you're forced into a slow and lossy compaction step. This was a larger consideration with our earlier reasoning models that had shorter context windows (~200k), longer thinking times (up to ~100k per message), and poor compaction. However, now that we've shipped longer context windows, we've trained our models think much more efficiently, and we've made compaction way better than it used to be, the balance of factors is changing. Tune in Thursday!
This is something I never understood. Why the reasoning is not included until the context is full, then the reasoning stripped optionally to allow the conversation to continue. and only then when its truly full offer a compaction. Was it to optimize caching? Well I guess it doesn't matter now that you hinted that this choice was made because of prior limitations and may change very soon
Models are typically trained (at longer conversations/more turns) either with or without the reasoning still in the conversation. If you train a model with those, then using it without them, the model will perform a lot worse, same vice-versa if you train without but then end up using the model with them.
That's why you'll see some models have it and others don't, and trying to use them another way, will make them worse, they weren't trained like that.
So why aren't the models trained with both? I'm guessing that sort of permutation in the training would lead to double the amount of training time being needed, as you know effectively will have two variants of every session you train on, with and without the reasoning.
Why do they store an encrypted reasoning payload in the session file and pass it to the API? Just a ruse? Reasoning isn’t even that many tokens, you think they’d degrade their model quality like that?
Reasoning messages would be lost immediately after a single tool call, unless you mean they sometimes go back and strip the reasoning channel retroactively, but that would increase costs via cache invalidation. I just don’t see any way it would make sense for them to do.
And wouldn’t this be noticeable by reasoning tokens not being accounted for in the context window usage?
Were you able to try Sol Ultra?
I think if you are not seeing reasonable performance in your agent loops as of 5.5, it's likely there is a deficit with how the loop, prompt or tools interact with the environment.
In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?
I would not be surprised if it is not as intelligent as the Mythos class models.
I have seen rumors that GPT 6 may release before September. The same person also claimed that a Fable 5.1 checkpoint has been completed a few weeks ago.
My quota is about to reset. Really can’t wait to use it.
So claude: 10 paragraphs of prose
codex: 1 paragraph of jargon over jargon.
UX is nicer where the agent is somehow "separated" from execution.
- alpha testers will start getting access now
- everyone will get access Thursday (barring banned countries / individuals)
Historically, some companies and individuals have gotten alpha access before public launches, to give feedback and adapt their products to the new models. With GPT-5.6, some folks had early alpha access, but this was paused while the model was being evaluated and approved. Now, alpha access will be enabled for partners in the next two days before our wider launch.
(I work at OpenAI.)