Anthropic Reveals Internal Model 2: Slightly More Capable Overall Than Claude Mythos 5
Anthropic has quietly disclosed an unreleased internal AI model called Model 2 in its August 2026 Risk Report . The model is not a public Claude release, and Anthropic has not anno

Anthropic Reveals Internal Model 2: Slightly More Capable Overall Than Claude Mythos 5
Introduction
Anthropic has quietly disclosed an unreleased internal AI model called Model 2 in its August 2026 Risk Report.
The model is not a public Claude release, and Anthropic has not announced it as a product. It appears in the report because the company includes internally deployed frontier and near-frontier systems when assessing risks from advanced AI.
According to Anthropic, Model 2 is somewhat more capable than Claude Mythos 5 overall.
The improvement is real, but the company is careful not to describe it as a major generational jump.
Anthropic’s own summary says Model 2 is noticeably better than Mythos 5 on many tasks relevant to internal use, while still falling well short of the kind of capability leap the company observed when moving from Claude Opus 4.6 to Claude Mythos Preview.
That makes Model 2 interesting for a different reason.
Rather than signaling a dramatic new generation, it looks more like a continued refinement of Anthropic’s frontier-model capabilities.
The report also gives one concrete data point: on Anthropic’s internal CoBench evaluation, Model 2 scored 62.8%, ahead of Claude Mythos 5 at 50.3% and Claude Mythos Preview at 54.8%.

The same report also mentions another internal system, Model 1, which has led to speculation about whether Anthropic is developing a sequence of internal models before assigning public Claude product names.
The official report supports only part of that interpretation.
Anthropic confirms that Model 1 and Model 2 are internal models, but it does not say that Model 2 is the formal successor to Model 1, nor that either will become a future Claude release.
Model 2 Is Slightly Stronger Overall Than Claude Mythos 5
The central point in the AIBase report is accurate: Anthropic says Model 2 is stronger overall than Claude Mythos 5.
The company’s August Risk Report describes Model 2 as:
- More capable than Mythos 5 in some areas.
- Less capable than Mythos 5 in others.
- Slightly more capable overall.
Anthropic also says its rough qualitative assessment is that Model 2 provides a noticeable improvement on many tasks that matter for internal use.
That wording is important.
Model 2 should not be described as universally better than Mythos 5 across every benchmark or every risk-relevant domain.
The report itself shows that the comparison varies by task.
For example, in some chemical and biological risk evaluations, Anthropic says the limited Model 2 results were comparable to or weaker than Claude Mythos 5.
In a separate covert-behavior evaluation, Model 2 was slightly stronger than Mythos 5 but still significantly weaker than Mythos Preview.
So the fairest summary is:
Model 2
= somewhat stronger overall
≠ better on every evaluation
The Improvement Is Not a Major Generational Leap
The original article emphasizes that the gain is relatively modest.
Anthropic says much the same thing.
The company specifically contrasts Model 2 with the jump from Claude Opus 4.6 to Claude Mythos Preview.
That earlier transition represented a much larger capability increase.
Model 2 does not.
Anthropic says the newer internal model is a noticeable improvement for many internal tasks, but not a discontinuity of the same scale.
This is also reflected in another internal capability indicator.
Anthropic maintains an internal version of the Epoch Capability Index called AECI, which combines performance across multiple internal benchmarks.
Based on limited data, Model 2 appears to score approximately:
1.5 AECI points
above Claude Mythos 5.
Anthropic explicitly notes that the estimate has large error bars.
It also says the increase is smaller than the improvement from Mythos Preview to Mythos 5.
That supports the article’s main characterization: Model 2 looks like an incremental frontier improvement rather than an obvious new capability regime.
CoBench Gives the Clearest Public Comparison
The most concrete evidence in the report comes from CoBench.
Anthropic describes CoBench as an internal evaluation designed to test models on actual research-and-development problems from the company rather than generic benchmark proxies.
The evaluation places a model at a historical point in Anthropic’s infrastructure.
The model receives a snapshot of resources such as:
- Code.
- Logs.
- Internal documentation.
- Historical internal messages.
It then has to diagnose the root cause of technical problems that Anthropic engineers previously solved.
The current version described in the August report contains:
449 problems
drawn primarily from issues solved between February and April 2026.
These problems come from parts of Anthropic’s technical organization involved in areas such as:
- Training models.
- Running model infrastructure.
- Debugging technical systems.
- Research engineering.
This makes CoBench especially relevant to Anthropic’s concern about AI automating AI research and development.
Model 2 Scores 62.8% on CoBench
The comparison shown in the report gives the following scores:
| Model | CoBench Score |
|---|---|
| Claude Sonnet 4.6 | 12.0% |
| Claude Opus 4.6 | 15.6% |
| Claude Opus 4.7 | 27.4% |
| Claude Mythos 5 | 50.3% |
| Claude Mythos Preview | 54.8% |
| Model 2 | 62.8% |
Model 2 therefore leads the models shown in the chart.
Its advantage over Mythos 5 is:
62.8% - 50.3% = 12.5 percentage points
That is meaningful on this particular internal benchmark.
It is still far below the score Anthropic believes would indicate full substitution for its technical staff.
Anthropic Estimates Full Researcher Substitution at Around 85%
Anthropic says that, based on manual inspection of the tasks and grading process, a system genuinely capable of fully substituting for Anthropic research scientists and engineers should score at least approximately:
85%
on CoBench.
Model 2’s 62.8% therefore leaves a substantial gap.
Anthropic also cautions that the benchmark is not a perfect measurement of human equivalence.
The dataset is deliberately filtered toward difficult problems.
Most tasks were selected because Mythos Preview failed them at least once in three attempts.
Without that filtering, Anthropic says the dataset would be roughly twice as large.
There are also other limitations:
- Some historical human conclusions may have been imperfect.
- Some answers may be ambiguous.
- Automated grading can make errors.
- The model does not receive every permission or infrastructure affordance a real Anthropic engineer might have.
So CoBench is best treated as a useful internal signal, not a universal measure of whether an AI model can replace a researcher.
Giving Mythos 5 More Tokens Does Not Close the Gap Completely
Anthropic also tested whether the CoBench result might simply reflect insufficient inference budget.
The figure uses a token budget of approximately:
300,000 tokens
for the evaluated setup.
When Anthropic increased Mythos 5’s budget to:
900,000 tokens
its score improved by only around:
3 percentage points
That suggests at least part of Model 2’s lead cannot be explained by simply giving Mythos 5 more tokens.
Anthropic still believes better scaffolding and evaluation harnesses could improve results, so the benchmark is not a pure measure of model weights alone.
Agent architecture matters.
Model 2 Is Already Used Heavily Inside Anthropic
Model 2 is unreleased, but it is not merely a laboratory checkpoint that nobody uses.
Anthropic says Model 2 and Claude Mythos 5 are among its most capable and most commonly used internal models.
They are used heavily for:
- Coding.
- Data generation.
- Other agentic workflows.
- Research.
- Engineering.
Anthropic also says both models are used through persistent agent deployments, not only interactive chat sessions.
This is significant because the company’s Risk Report is partly concerned with the possibility that AI systems could automate a growing portion of Anthropic’s own R&D.
The report says Claude now writes a large majority of the code merged into Anthropic’s production codebases.
However, Anthropic does not conclude that its models can yet fully replace its researchers.
The company believes AI is materially accelerating internal work, but its August assessment remains below the Responsible Scaling Policy threshold for dramatic AI-driven R&D acceleration.
Model 2 Has Not Completed Anthropic’s Full Predeployment Evaluation Suite
The report contains another important caveat.
Model 2 has gone through Anthropic’s internal deployment review process, but it has not undergone the full suite of assessments the company would normally run before an external deployment.
Anthropic therefore says it has somewhat lower confidence in its understanding of Model 2 than it does for a fully evaluated release model.
That matters when interpreting benchmark comparisons.
A single CoBench result can show that Model 2 is strong at internal R&D tasks.
It does not establish a complete performance profile across:
- General reasoning.
- Coding.
- Cybersecurity.
- Biology.
- Safety.
- Alignment.
- Long-horizon autonomy.
- Consumer use.
- Enterprise workloads.
Anthropic’s current conclusion is deliberately broad: Model 2 appears slightly more capable overall than Mythos 5.
Anthropic Currently Has No Plans to Release Model 2 Externally
The most important practical detail is easy to miss in the short AIBase article.
Anthropic explicitly says:
We do not currently have plans
to release this model externally.
So Model 2 should not be treated as an upcoming Claude product with an implied launch date.
There is currently no official:
- Release date.
- Public model name.
- Claude API model ID.
- Pricing.
- Context-window specification.
- Consumer availability plan.
- Enterprise rollout plan.
Any claim that Model 2 will become “Claude Mythos 6,” “Claude Opus 6,” or another named product would be speculation.
Anthropic disclosed the system because its internal models are relevant to the company’s risk assessment.
That is different from announcing a product.
Claude Mythos 5 Is Not Actually a Generally Available Public Model
The original AIBase headline refers to Claude Mythos 5 as Anthropic’s strongest publicly available model.
That needs a small correction.
Claude Mythos 5 is not generally available.
Anthropic currently offers it to a limited group of approved customers through Project Glasswing and related trusted-access programs.
The generally available counterpart is Claude Fable 5.
Anthropic says Fable 5 and Mythos 5 use the same underlying model.
The difference is that Fable 5 includes stronger safeguards for sensitive areas such as cybersecurity and biology, while Mythos 5 is available to vetted users who need fewer restrictions for authorized work.
The current product relationship is:
Claude Mythos 5
= limited trusted access
Claude Fable 5
= same underlying model
+ stronger safeguards
+ broad availability
For a general reader, it is therefore more precise to call Mythos 5 Anthropic’s strongest limited-access frontier model, while Fable 5 is its most capable widely released model.
Model 1 Is Real, but the “Model 1 → Model 2” Product Story Is Speculation
The report also mentions Model 1.
That is enough to naturally invite speculation:
Model 1
→ Model 2
→ future Claude release?
But Anthropic does not make that claim.
Its description of Model 1 is actually fairly restrained.
The company says Model 1 is broadly similar in capability to:
- Claude Mythos Preview.
- Claude Mythos 5.
It received relatively limited internal deployment.
Anthropic says a strong majority of internal users preferred Mythos 5 and that Model 1 usage was low and declining by the report’s July 15 coverage date.
The company also says:
it does not expect Model 1
to be externally deployed
or widely used internally in the future
So Model 1 appears more like an internal development branch or checkpoint than a clearly defined previous generation waiting for public release.
Model 2 Is Different From Model 1 in One Important Way
Anthropic’s treatment of Model 2 is noticeably more serious.
Model 1 is discussed briefly and then largely excluded from the risk analysis because Anthropic considers Mythos 5 a reasonable upper bound on its risks.
Model 2, by contrast, is used throughout major parts of the report because it is:
- More capable overall.
- More heavily used internally.
- Relevant to Anthropic’s frontier-risk analysis.
The report says Model 2’s internal usage is similar to Mythos 5.
That does not prove a public release is coming.
It does show that Model 2 is operationally important inside Anthropic.
The Risk Report Is Why the Model Became Public Knowledge
Model 2 was not introduced through a conventional launch blog post.
It appeared because Anthropic’s Responsible Scaling Policy requires the company to analyze risk from systems it is actually developing and deploying internally.
The August 2026 Risk Report covers Anthropic’s activities through a July 15, 2026 coverage date.
Anthropic lists internal models when they are relevant to questions such as:
- Misalignment in high-stakes settings.
- Automated AI R&D.
- Biological and chemical risk.
- Model security.
- Internal monitoring.
That is why the report contains more information about Model 2 than a normal unreleased model would receive.
The disclosure is a safety-governance artifact first and a product leak second.
What the Report Does and Does Not Tell Us
| Claim | Status |
|---|---|
| Anthropic has an internal model called Model 2 | Confirmed |
| Model 2 is unreleased | Confirmed |
| Model 2 is slightly more capable overall than Mythos 5 | Confirmed by Anthropic |
| Model 2 is better than Mythos 5 in every area | No |
| Model 2 scores 62.8% on the CoBench comparison shown in the report | Confirmed in the reported figure |
| Mythos 5 scores 50.3% on the same comparison | Confirmed in the reported figure |
| Anthropic estimates full technical-staff substitution at roughly 85% on CoBench | Confirmed |
| Model 2 is heavily used internally | Confirmed |
| Model 2 has completed the full normal external-release evaluation suite | No |
| Anthropic currently plans to release Model 2 externally | No |
| Model 2 has an official Claude product name | No |
| Model 1 is a real internal model | Confirmed |
| Model 1 is definitely the direct predecessor of Model 2 | Not established |
| Model 1 is expected to launch publicly | No |
| Claude Mythos 5 is generally available to everyone | No |
| Claude Fable 5 is the broadly released counterpart to Mythos 5 | Confirmed |
常见问题
What is Anthropic Model 2?
Model 2 is an unreleased internal frontier model disclosed in Anthropic’s August 2026 Risk Report. Anthropic says it is slightly more capable overall than Claude Mythos 5 and is already used heavily for coding, data generation, research, engineering, and other agentic work inside the company.
Is Model 2 better than Claude Mythos 5?
Overall, Anthropic says yes—but only slightly. The company also says Model 2 is stronger in some areas and weaker in others, so it should not be described as universally superior.
What is Model 2’s CoBench score?
The comparison shown in the report gives Model 2 a score of 62.8%, ahead of Mythos Preview at 54.8% and Mythos 5 at 50.3%. Anthropic estimates that a model fully capable of substituting for its technical staff would need to reach at least roughly 85% on the current evaluation.
What does CoBench measure?
CoBench is Anthropic’s internal evaluation based on historical research-and-engineering problems that its own technical staff previously solved. Models receive historical snapshots of code, logs, internal messages, and documentation and must diagnose the underlying technical issue.
Will Anthropic release Model 2?
Anthropic says it currently has no plans to release Model 2 externally. The company has not announced a product name, release date, API identifier, pricing, or public-access plan.
Is Model 2 the next Claude Mythos model?
Anthropic has not said that. Model 2 is an internal identifier used in the Risk Report, and assigning it a future Claude product name would be speculation.
What is Anthropic Model 1?
Model 1 is another internal model disclosed in the same report. Anthropic says it is broadly similar to Mythos Preview and Mythos 5, has relatively low and declining internal usage, and is not expected to be externally deployed.
Is Claude Mythos 5 publicly available?
Not generally. Mythos 5 is available through limited trusted-access programs such as Project Glasswing, while Claude Fable 5 uses the same underlying model with stronger safeguards and is broadly available.
相关工具
- Claude: Anthropic’s consumer and professional AI assistant for general-purpose work.
- Claude API: Anthropic’s developer platform for generally available Claude models.
- Anthropic Console: The official workspace for testing Claude API models and managing developer access.
- Project Glasswing: Anthropic’s trusted-access initiative for advanced Mythos-class cybersecurity capabilities.
- Anthropic Transparency Hub: Central resource for model safety evaluations, system cards, and risk information.
Related Links
- Anthropic August 2026 Risk Report: The primary source disclosing Model 1, Model 2, CoBench, internal usage, and current release plans.
- Claude Fable 5 and Claude Mythos 5: Anthropic’s official launch announcement explaining the relationship between the two models.
- Claude Models Overview: Current official availability and API information for Anthropic’s released models.
- Claude Mythos: Official information about Mythos 5 availability, safeguards, pricing, and Project Glasswing.
- Anthropic Responsible Scaling Policy: The governance framework behind Anthropic’s periodic risk-reporting process.
- Anthropic Transparency Hub: Official safety and capability information for current Claude models.
Summary
Anthropic’s August 2026 Risk Report reveals an internal model called Model 2 that is already being used heavily for research, engineering, coding, data generation, and persistent agent workflows.
The model is stronger than Claude Mythos 5 on Anthropic’s internal CoBench evaluation, scoring 62.8% versus 50.3%. Anthropic also says Model 2 appears slightly more capable overall, although it is weaker than Mythos 5 in some areas and the improvement is much smaller than earlier major capability jumps.
The report also mentions Model 1, but the evidence does not support treating Model 1 and Model 2 as a confirmed public product sequence. Model 1 is seeing declining internal use, while Anthropic explicitly says it currently has no plans to release Model 2 externally.
The clearest takeaway is that Anthropic’s internal frontier has already moved slightly beyond Mythos 5—but Model 2 is an internal research system, not an announced next-generation Claude product.