NEX-N2.5 / OVERVIEW

在真实环境中持续行动Sustained Action in Real Environments

Nex-N2.5 包含 Mini、Pro 和 Max 三种尺寸。Nex-N2.5 comes in three sizes: Mini, Pro, and Max.

模型系列The Model Family

Mini 与 Pro 延续 Nex-N2 的多模态底座,重点强化计算机使用、网页浏览与视觉辅助的能动性,让视觉理解与环境交互共同服务于真实任务。Max 基于 DeepSeek-V4-Pro-Base,采用 1.6T 参数的 MoE 纯文本架构。我们首次在万亿参数规模上开展完整后训练,探索更大模型在复杂 Agent 任务中的能力。Mini and Pro build on Nex-N2’s multimodal base, improving computer use, web browsing, and visually assisted agency to connect visual understanding with real-world action. Max builds on DeepSeek-V4-Pro-Base, a 1.6-trillion-parameter, text-only MoE base. Our first full post-training effort at this scale explores larger models’ capabilities on complex agent tasks.

视觉反馈Visual Feedback

面向真实环境中的长程任务,Nex-N2.5 依靠视觉反馈持续行动与自我修正,操作计算机、浏览器并运行和测试程序。视觉成为 Agent 感知环境、验证结果与推进任务的重要接口,让模型能够依据环境变化调整后续行动。For real-world, long-horizon tasks, Nex-N2.5 uses visual feedback to act and self-correct, operate computers and browsers, and run and test code. Vision serves as an interface for perceiving the environment, verifying results, and advancing tasks, helping agents adapt their next actions to changes they observe.

环境与训练Environments and Training

我们进一步扩充 Agent 训练环境、任务类型与生产力场景,结合多样化任务和环境反馈,开展系统性后训练。训练强化了模型在科学研究、知识工作和复杂生产力任务中的能力,也为更大规模模型的能动性训练积累实践经验。We expanded agent training environments, task types, and productivity scenarios, using diverse tasks and environmental feedback for systematic post-training. This strengthens scientific research, knowledge work, and complex productivity capabilities, while building experience in agency training for larger models.

真实生产力场景评测Real-World Productivity Benchmarks

Nex-N2.5 Mini、Pro、Max 在编程、工具使用和多模态任务中的表现。Nex-N2.5 Mini, Pro, and Max across coding, tool use, and multimodal tasks.

Text Benchmarks

Multimodal Benchmarks

用 Nex-N2.5 构建Built with Nex-N2.5

使用 NexRT 推理Inference with NexRT

NexRT 是面向 Nex-N2.5-Pro 的单请求推理引擎,支持在 8 张 H100 或 H200 GPU 上运行普通解码、MTP 和 DFlash。通过定制 CUDA 内核、完整解码 CUDA Graph 和 GPU 直接通信,NexRT 优化从嵌入、注意力、MoE 到采样的整个解码路径。NexRT is a specialized single-request inference engine for Nex-N2.5-Pro on 8 H100 or H200 GPUs. It supports standard decoding, MTP, and DFlash, combining custom CUDA kernels, complete decode graphs, and direct GPU communication to optimize the path from embedding through attention and MoE to sampling.

NexRT 使用 SGLang 完成预填充,并在预填充与解码引擎之间共享路由专家权重;上下文并行注意力进一步减少长上下文中的重复 KV 读取。SGLang handles prefill, with routed MoE weights shared between the prefill and decode engines. Context-parallel attention reduces repeated KV reads at long contexts.

NexRT generation demo
Nex-N2.5-Pro 文本生成演示。点击查看原始动图。Text generation with Nex-N2.5-Pro. Open the original animation for a closer look.
Nex-N2.5-Pro decode throughput on 8×H100 with and without CP
8×H100 · TP=8。MTP 与 DFlash 吞吐量由完整 CUDA Graph 轮次耗时计算,假设平均每轮提交的 token 数(AL)分别为 3 和 5.5;实际吞吐量取决于草稿接受率。CP 表示上下文并行。8×H100 · TP=8. MTP and DFlash rates are calculated from complete CUDA-graph rounds, assuming average committed lengths (AL) of 3 and 5.5 respectively; actual throughput depends on draft acceptance. CP denotes context parallelism.

评估Evaluation

Notes:

Score sources : Where available, scores are drawn from official benchmark leaderboards and the latest evaluation reports published by model providers, including the Kimi-K3, Qwen3.8-Max, GLM-5.3, and HY4 reports. Results without a public source are obtained through our own evaluations.

Sampling parameters : Our evaluations use temperature = 0.7, top_p = 0.95, and top_k = 40.

Evaluation harness : Coding tasks are evaluated using the NexAU harness.

DeepSeek-V4-Pro : Our evaluations use the DeepSeek-V4-Pro-0813 version.

BrowseComp : We apply the Summary context-compaction strategy when the token usage exceeds 60% of the model’s context window.

Vision2Web : We report the average score across the Frontend, Webpage, and Website categories, with Gemini-3.5-Flash as the VLM judge and GLM-5V-Turbo (Claude Code) as the GUI agent.

Computer-use and browser-use benchmarks, including OSWorld, WebTest, and WebArena, are evaluated using our NexCUA harness. Grounding coordinates are normalized to a 0–1000 scale. The NexCUA project will be open-sourced soon.

WebTestBench : These results are evaluated in oracle mode, using the ground-truth checklist to assess defect detection only, without checklist generation.

Notation : Bold marks the best result in each benchmark, including ties; — indicates unavailable data.

模型开源Open-Source Models

Nex-N2.5-Mini

基座 · Qwen3.5-35B-A3B-BaseBase · Qwen3.5-35B-A3B-Base

适合高速指令遵循、实时工具执行以及高性价比规模化部署。Suited for high-speed instruction following, real-time tool execution, and cost-effective large-scale deployment.

参数Parameters
35B (A3B)
架构Architecture
MoE

Nex-N2.5-Pro

基座 · Qwen3.5-397B-A17BBase · Qwen3.5-397B-A17B

适合复杂推理、多智能体编排以及高级软件工程任务。Suited for complex reasoning, multi-agent orchestration, and advanced software engineering tasks.

参数Parameters
397B (A17B)
架构Architecture
MoE

Nex-N2.5-Max

基座 · DeepSeek-V4-Pro-BaseBase · DeepSeek-V4-Pro-Base

面向科研建模、工程设计与复杂前端开发等专业任务。For scientific modeling, engineering design, and complex frontend development.

参数Parameters
1.6T (A49B)
架构Architecture
MoE