MaHui|13-Hour Nonstop Coding: Moonshot AI's K2.6 Reconstructs Open-Source Engine with 185% Performance Surge


Recently, MaHui member Moonshot AI released and open-sourced the Kimi K2.6 model, delivering industry-leading, state-of-the-art capabilities in coding, long-horizon task execution, and agent swarms.
Kimi K2.6 is now live on kimi.com, the latest version of the Kimi app, the Kimi API, and Kimi Code — the programming assistant. All users can start using it now.

See the technical blog for full benchmark results. Kimi K2.6 has achieved comprehensive improvements in general agent capabilities, coding, and visual understanding. It posted industry-leading scores on benchmarks including the full, PhD-level Humanity's Last Exam, SWE-Bench Pro (which tests real-world software engineering ability), and DeepSearchQA (which evaluates agent deep-retrieval capability) — matching or outperforming closed-source models such as GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
Kimi K2.6 is Kimi's strongest coding model to date. Its long-horizon coding capability has also improved significantly. In testing, it can code uninterrupted for 13 hours, writing or modifying more than 4,000 lines of code to complete the development and optimization of complex systems. By deeply integrating code with visual capabilities, K2.6 has elevated code-driven design to a new level, capable of delivering professional-grade web applications with striking creative design.
Kimi K2.6 substantially enhances agent autonomous execution, helping Kimi further expand the scope of agent capabilities:
- The "Agent Swarm" architecture powered by the K2.6 model has received a major upgrade. It now supports 300 sub-agents working in parallel across 4,000 collaborative steps, achieving greater scale in parallelization while significantly improving task completion and delivery quality compared to K2.5.
- For proactive agent frameworks such as OpenClaw and Hermes Agent, K2.6 demonstrates exceptional automated task processing, supporting up to 5 days of continuous autonomous operation.
Breakthrough in Long-Horizon Coding
K2.6 has achieved a breakthrough in long-horizon coding tasks, demonstrating more reliable generalization across different programming languages (such as Rust, Go, Python) and task scenarios (such as frontend, DevOps, and performance optimization).
On Kimi Code Bench — Kimi's internal, rigorous code evaluation benchmark covering a variety of complex end-to-end tasks — K2.6 improved approximately 20% over K2.5.

According to Moonshot AI's real-world testing, the Kimi K2.6 model demonstrated powerful long-horizon reasoning in complex software engineering tasks:
Scenario 1: K2.6 successfully downloaded and deployed the Qwen3.5-0.8B model locally on a Mac, implementing and optimizing model inference using the niche Zig language — proving the new model's generalization ability. After more than 4,000 tool calls and over 12 hours of uninterrupted operation, the K2.6 model iterated through 14 rounds, boosting throughput from roughly 15 tokens/s to roughly 193 tokens/s, ultimately achieving inference speed 20% faster than LM Studio.

Scenario 2: Kimi K2.6 autonomously completed a deep refactoring of exchange-core, an open-source financial matching engine with 8 years of history. Over 13 hours of continuous work, the model iterated through 12 optimization strategies, making precise modifications to more than 4,000 lines of code through over 1,000 tool calls. Acting as an expert systems architect, Kimi K2.6 analyzed CPU and memory allocation flame graphs to pinpoint hidden bottlenecks, and boldly restructured the core thread topology (from 4ME+2RE to 2ME+1RE). Even though the engine's performance was already near its limit, Kimi K2.6 still achieved a 185% median throughput jump (from 0.43 to 1.24 MT/s), with peak throughput surging 133% (from 1.23 to 2.86 MT/s).

A New Benchmark for Code-Driven Design
Beauty itself is a form of productivity. K2.6 Agent mode can now produce websites with exceptional design sensibility and visual impact.
With proficient use of image and video generation tools, the K2.6 Agent can generate visually cohesive assets, construct hero sections with strong visual focal points, and implement interactive elements along with rich scroll-triggered animations.
The K2.6 Agent isn't limited to writing frontend pages — it also supports basic backend database modules, such as embedding form-based information collection within generated web pages.
With stronger multimodal programming capabilities, K2.6 can more accurately translate image and video assets into code:
Moonshot AI created a dedicated frontend development and design evaluation benchmark (Kimi Design Bench), covering four dimensions: visual input tasks, landing page construction, full-stack application development, and general web development. Compared to the Gemini 3 model in Google AI Studio, the K2.6 Agent based on the kimi.com framework demonstrated a very clear lead.

Agent Swarm Fully Upgraded
Breaking through the performance limits of single agents is the only way to scale agent capabilities. "Agent Swarm" is a new capability Moonshot AI introduced starting with the K2.5 model — dynamically decomposing complex tasks and autonomously spawning specialized agents to process them in parallel.

Building on K2.5, K2.6's Agent Swarm collaboration has been comprehensively upgraded. The Agent Swarm can now dispatch agents with different skill specialties to complement each other, combining capabilities in search, deep research, document analysis, and long-form writing — with significantly improved task completion quality compared to K2.5. In a single run, the Agent Swarm can independently deliver end-to-end multi-product outputs, from documents to web pages to PPTs and spreadsheets.
The Agent Swarm architecture has also been upgraded, now supporting up to 300 sub-agents working in parallel across 4,000 collaborative steps, achieving greater parallelization and further pushing the ceiling of multi-agent system collaboration.
Two use cases:
Case 1: The Agent Swarm designed and executed 5 quantitative strategies for 100 global semiconductor stocks. It distilled McKinsey-style PPT logic into reusable skills, ultimately delivering detailed modeling spreadsheets and a complete set of presentation documents.
Case 2: The Agent Swarm transformed a high-quality astrophysics paper containing massive amounts of visual data into a reusable academic skill. By extracting the paper's reasoning flow and visualization methods, the system produced a 40-page, 7,000-word research paper, along with a structured dataset containing over 20,000 entries and 14 publication-grade astronomical charts.
Autonomous Agent: Seamless Integration with OpenClaw/Hermes and Other Frameworks
K2.6 significantly enhances agent autonomous execution, particularly excelling in OpenClaw and Hermes Agent-style automation — scenarios that require AI to operate 24/7 across applications.
Unlike traditional conversational interaction, these workflows require AI to actively manage task planning, execute code, and coordinate cross-platform operations as a persistent background agent.
Moonshot AI's RL infrastructure team used a K2.6-based agent to achieve 5 consecutive days of autonomous operation. The agent handled monitoring, incident response, and system operations, demonstrating sustained context maintenance, multi-threaded task processing, and full-cycle execution from alert receipt to complete resolution. Below is the K2.6 work log (sensitive information anonymized):
K2.6 has seen tangible reliability improvements in real-world use: more precise API calls, more stable long-duration operation, and enhanced safety awareness when executing complex research tasks.

Moonshot AI's internal Claw Bench test results show that K2.6 improved 10% over K2.5 in overall performance. This benchmark covers five dimensions: programming tasks, instant messaging ecosystem integration, information retrieval and analysis, scheduled task management, and memory recall. Across all evaluation metrics, K2.6 leads K2.5 in both task completion rate and tool call accuracy, with particularly significant advantages in workflows that require long-duration autonomous operation without human intervention.
Office Productivity Continues to Improve

Leveraging K2.6's stronger coding and visual understanding capabilities, Kimi Agent mode now supports creating and invoking Skills.
The system comes with over 100 officially recommended skills built in. These include an investment research skill pack created by Moonshot AI's internal expert team, which packages institutional-grade investment research workflows to let users generate professionally formatted one-pagers or in-depth research reports on A-share, Hong Kong, and US-listed companies with a single click — quickly getting up to speed on a company's key fundamentals, industry landscape, and the core stock price drivers the market cares about most.
Moonshot AI will continue updating the recommended skill library to help more knowledge workers achieve "plug-and-play" efficiency gains across the full workflow from finding materials, organizing ideas, to producing deliverables.
Starting now, type a slash "/" in Kimi Agent mode to begin creating and invoking skills. Every user can create skills from scratch through conversation with Kimi.

But creating truly practical skills still requires substantial knowledge and professional expertise — the barrier remains high. To help users easily transform their carefully crafted documents into reusable Skills, Kimi Agent now supports "Office Document to Skill": upload a high-quality Office document, and Kimi will attempt to understand the original document's structure and stylistic DNA to generate a bespoke, reusable document creation skill for you.

One More Thing
Through teamwork and organizational division of labor, humanity created the internet, built large models, and landed on the moon. If AI agents are to help humans tackle complex real-world problems, they too must evolve toward teamwork and organizational division of labor.
"Agent Swarm" is Moonshot AI's exploration in the direction of AI automated division of labor. Today, the company begins exploring another direction: putting humans and various 24/7 agents together in a group — how can they divide labor and collaborate to accomplish tasks that neither a single person nor a single agent could complete?

This is "Claw Groups" — now in limited beta.
"Claw Groups" aims to embrace an open, heterogeneous ecosystem: multiple agents and humans operating as true collaborators. Users can connect 24/7 agents from any device, any vendor, running any model (initially supporting OpenClaw, with Hermes Agent and other frameworks to follow). Each agent can bring its own professional toolkit, skills, and persistent memory context. Whether deployed on a local laptop, mobile device, or cloud instance, these diverse agents can join the same collaborative workspace.
In "Claw Groups," K2.6 serves as the coordinator. It dynamically matches tasks to agents based on their skill profiles and available tools, achieving optimal capability allocation. When an agent encounters a failure or stalls, the coordinator detects the interruption, automatically reassigns tasks or spawns sub-tasks, and actively manages the full lifecycle of agent deliverables from initiation through verification to completion.
Kimi Claw users will gradually receive invitations to the "Claw Groups" beta. Stay tuned.
Start Using Kimi K2.6
Kimi K2.6 is now available to all free users, paid subscribers, Kimi Code users, and enterprise API users. Visit kimi.com, the latest Kimi App, Kimi Code, and the Kimi API Open Platform (platform.kimi.com) to get started.
Enterprises and developers can begin using it immediately by specifying kimi-k2.6 as the model in the Kimi API. To celebrate the K2.6 model API launch, the Kimi Open Platform is running a limited-time bonus offer of up to 30% extra credit on deposits.

Meanwhile, the official Kimi K2.6 API has debuted on Tencent Cloud TokenHub and other platforms. Tencent Cloud users are welcome to try out the Kimi K2.6 model. Additionally, Moonshot AI recommends calling the official Kimi API directly to reproduce Kimi K2.6 benchmark results. For those using third-party API services, the Kimi Vendor Verifier (KVV) can help identify higher-precision service providers. Learn more: https://kimi.com/blog/kimi-vendor-verifier



