Notes:
Score sources : Where available, scores are drawn from official benchmark leaderboards and the latest evaluation reports published by model providers, including the Kimi-K3, Qwen3.8-Max, GLM-5.3, and HY4 reports. Results without a public source are obtained through our own evaluations.
Sampling parameters : Our evaluations use temperature = 0.7, top_p = 0.95, and top_k = 40.
Evaluation harness : Coding tasks are evaluated using the NexAU harness.
DeepSeek-V4-Pro : Our evaluations use the DeepSeek-V4-Pro-0813 version.
BrowseComp : We apply the Summary context-compaction strategy when the token usage exceeds 60% of the model’s context window.
Vision2Web : We report the average score across the Frontend, Webpage, and Website categories, with Gemini-3.5-Flash as the VLM judge and GLM-5V-Turbo (Claude Code) as the GUI agent.
Computer-use and browser-use benchmarks, including OSWorld, WebTest, and WebArena, are evaluated using our NexCUA harness. Grounding coordinates are normalized to a 0–1000 scale. The NexCUA project will be open-sourced soon.
WebTestBench : These results are evaluated in oracle mode, using the ground-truth checklist to assess defect detection only, without checklist generation.
Notation : Bold marks the best result in each benchmark, including ties; — indicates unavailable data.