- https://www.longcatai.org/models/longcat-2
- 1.6T MoE (activates ~33–56B, ~48B avg) trained for agentic coding. Native 1M context via LongCat Sparse Attention (linear, not quadratic).
- Zero-computation experts skip easy tokens; MOPD fuses agent / reasoning / interaction expert groups into one model. Trained on 30T+ tokens on a 50k-card domestic cluster.
- Headline numbers they cite: SWE-bench Pro 59.5, Terminal-Bench 2.1 70.8, BrowseComp 79.9. Preview on OpenRouter / longcat.ai; weights still pending HF.
Connections
No outgoing connections.