2026 MaHui | DeepWise's Weijie Sun: Building an AI Scientist That Surpasses Humans, the Anthropic of Scientific Research
DeepWise has been working on something difficult but worthwhile — building a shared scientific foundation for research and discovery worldwide.

At the 2026 MaHui summit, Weijie Sun, founder of DeepWise — an early portfolio company of Source Code Capital — delivered a keynote on "AI for Science." From its founding, DeepWise systematically began building the necessary infrastructure. By 2025, AI for Science had become a national-level priority in China, the United States, and Europe. Throughout, DeepWise has pursued something difficult but essential: constructing a shared scientific foundation for research and discovery worldwide.
The following is an edited selection of Sun's remarks:

Thank you to MaHui for the invitation. Every time I attend, I feel genuinely proud and warmly supported. DeepWise is a project Source Code invested in at an early stage. At our core, we do AI for Science.
The ultimate goal of AI for Science is to create AI scientists that comprehensively surpass human capability. Much as AI may fully exceed human coding ability this year or next, the core problem AI for Science addresses has now gained broad global recognition.
1
From "First to Propose" to "Global Consensus"
DeepWise was founded in 2018 as a pioneer in AI for Science. We began systematically building relevant infrastructure from day one. The early path of exploration was exceptionally difficult, and for that we are especially grateful to Source Code for its sustained support and companionship from the beginning. It wasn't until 2025 that AI for Science finally achieved global industry consensus.
In 2025, China, the United States, and Europe all introduced national-level policies: China's "AI Plus" action plan, released in July, listed AI for Science as its top priority; Europe's Horizon program, announced in August, ranked it second in strategic importance; and in November, the United States launched the Genesis program, elevating it to the level of fundamental national policy.
To put this in perspective, the United States has only launched two other programs of equivalent stature in its history: the Manhattan Project, which developed the atomic bomb, and the Apollo program, which secured America's Cold War technological advantage. Genesis is the third. The strategic intent is unmistakable.
Why does AI for Science command such high strategic status? In short, it is the core gateway through which all of humanity will acquire scientific knowledge and conduct research in the future, the underlying engine for technological innovation across every industry, and a critical chip in great-power technological competition. Notably, the six strategic dimensions outlined in the U.S. Genesis program align almost completely with the "Four Pillars and N Columns" infrastructure framework we proposed in 2022.

2
The Dilemma of Traditional Research, and the Answer from AI Scientists
So what underlying problem does AI for Science actually solve? Our answer: not merely how AI can achieve breakthroughs in one or two narrow domains, but how to create a shared scientific foundation for research and discovery worldwide.
Traditional human research has always operated at extremely low input-output ratios — far below factories, IDC facilities, or internet platforms. The numbers tell the story: global R&D spending totals roughly $2.8 trillion annually; China alone invests 3.6 trillion RMB per year, about 2.7% of GDP. Yet this massive investment yields remarkably limited scientific output.
The root cause lies in fundamental flaws in the traditional research paradigm. Conventional science centers on human scientists using tools to explore the world — essentially "read, compute, and do": reading literature, performing calculations, conducting experiments. Two critical pain points emerge here.
First, the tools are outdated and inefficient. To this day, our approach to reading literature differs little from 17th-century scholars flipping through printed papers. Research software and experimental instruments have failed to fully absorb the productivity gains of successive technological revolutions.
Second, talent cannot be replicated. As the subject of research, humans cannot duplicate, amplify, or compound their scientific expertise, leading directly to severe shortages of top talent. Of roughly 60 to 80 million active researchers worldwide, fewer than 2 million are elite scientists capable of core breakthroughs in new drug development, novel materials, or foundational science.
AI for Science will fundamentally change this. On one hand, it can intelligently upgrade and replace traditional research tools. On the other, AI agents can already simulate humans, independently taking on research tasks and completing full closed loops. This is DeepWise's core vision: building AI scientists that comprehensively surpass humans,彻底解决 the global shortage of top research talent.
Based on technological trajectories and our internal capabilities, we believe AI agents capable of fully surpassing human scientists could take shape within two to three years. The equally certain opportunity right now is the intelligent upgrading of traditional research infrastructure — which simultaneously equips AI scientists with knowledge, computing power, and experimental capabilities.
Simply boosting compute and execution cannot solve complex scientific problems; it requires a mature "read, compute, and do" tool system. Toward this goal, we have built a research closed loop centered on research agents with integrated "read, compute, and do" capabilities — omit any of these four elements, and the AI scientist cannot function.
We have already developed AI scientists with 5–8 years of equivalent research experience in multiple specialized domains. One example is MatMaster, an AI materials scientist co-developed with the Suzhou Laboratory of Materials Science. These AI scientists are built on our underlying general-purpose scientific intelligence foundation, SciMaster. Internal testing shows SciMaster performs at postdoctoral level across all disciplines, already surpassing human postdocs in core competencies including literature review, scientific computing, experimental operation, and manuscript writing.

3
Read, Compute, and Do: Three Infrastructure Systems and Commercial Implementation
Around "read, compute, and do," we have systematically reconstructed three major research infrastructure systems.
Knowledge system (Read). We have completed full-scale structured processing of massive volumes of high-quality global literature, patents, and data. AI agents can search all research knowledge resources and produce literature reviews with precise sourcing and minimal hallucination. A deep research report that takes 10 minutes to generate equals one month of work by core R&D personnel — a 10x efficiency gain. The challenge lies in data cleaning: manual cleaning and annotation of a single PDF is prohibitively expensive, and with hundreds of millions of documents globally, pure manual processing is impossible. Through human-AI collaborative annotation, we have compressed costs to roughly one-thousandth. Since the launch of Bohr·Science Navigation, registration and traffic have ranked first in the industry, with 4.35 million researchers worldwide currently using our products.
Computing system (Compute). Computational tools are the core foundation of research iteration. We have completed intelligent adaptation and MCP transformation of various research computing tools from open-source communities and public platforms, with over 50,000 categories of scientific computing software and tools now deployed on our platform. For domains where public tools cannot achieve high-precision calculations, we develop proprietary scientific foundation models, including: Uni-SMART for multimodal scientific literature, Uni-AIMS for representation, Uni-Fold for proteins, Uni-RNA for genes, DPA for atoms, and Uni-MOL for molecular conformations — meeting high-precision computing needs in materials, chemistry, biology, and other fields. Notably, in May 2026, DeepWise, as a core contributor, jointly launched the next-generation model architecture DPA4 with partners. On Matbench Discovery, the authoritative international benchmark for materials discovery, DPA4 achieved world-first ranking by the comprehensive performance metric CPS, becoming the new SOTA model. Simultaneously, on SPICE-MACE-OFF, the authoritative molecular benchmark, DPA4 achieved new SOTA results with fewer parameters, outperforming the previous leading model eSEN to take first place.
Experiment system (Do). All frontier discoveries ultimately require laboratory synthesis, testing, and validation. The current core pain point is the inability of agents to efficiently interface with physical laboratory hardware, lacking a unified underlying operating system. We were the first to enter this track, developing our proprietary UniLab OS intelligent laboratory operating system, which has completed integration and standardized management of over 150 major categories and nearly 2,000 types of experimental instruments and equipment — truly enabling AI to take over laboratories and conduct experiments autonomously.
The efficiency advantages of this system are dramatic. In polymer materials R&D, where a PhD student previously needed one morning to organize 100 reagent sets and calibrate equipment, our intelligent system achieves daily output equivalent to dozens of PhD students' full-day work. In Shanghai, our intelligent peptide drug molecule discovery platform alone matches the output of 500–800 chemists working around the clock; in Yibin, our electrolyte and solid-state electrolyte intelligent R&D system equals the capacity of nearly 100 specialized personnel.
With infrastructure in place, how to close the commercial loop? Our capabilities can all be packaged as standardized software services for customers; for large enterprises and research institutions, we provide customized R&D solutions. Here I want to highlight the FDE (Frontline Deployment Engineer) model, which is exceptionally well-suited to the research domain due to three industry pain points:
First, R&D demands are extremely fragmented and long-tailed. For a cathode materials client we serve, we mapped over 160 distinct R&D process steps end-to-end — no single-point tool can cover this; yet saving 30–40% time at each step compounds into disruptive efficiency gains across the full process.
Second, the field faces a triple barrier spanning AI technology, scientific domain expertise, and engineering implementation. Those who understand algorithms don't know materials chemistry; those with domain expertise lack AI literacy — making pure product-based solutions inadequate.
Third, R&D data is core corporate confidential information that companies will never share externally, rendering models dependent on customer data iteration unsustainable.
Based on these three factors, the optimal business model is: 70–80% standardized platform capabilities as the foundation, plus 20–30% customized development, deeply adapted to each enterprise's R&D workflow. To execute this, we have assembled "three-in-one FDE special forces teams" — members combining domain scientist, AI algorithm expert, and product delivery expert capabilities, capable of precisely translating customer needs into AI-deployable problems and delivering rapidly.
To summarize DeepWise's core advantages in three keywords: earliest, deepest, fastest. We were the first globally to define AI for Science and systematically build research infrastructure; we have the deepest investment in underlying core infrastructure with the strongest generality and fundamentality; and we lead the industry in coverage and commercial implementation speed. Currently, millions of researchers worldwide and over 100 domestic universities use our platform, with hundreds of leading enterprises in life sciences, physical sciences, and other fields in deep cooperation with us.
4
The Next Five Years and 2040: A Disruptive Transformation
The direction of research development aligns fully with China's 15th and 16th Five-Year Plan strategies. The core national goals are achieving basic technological self-reliance and resolving "chokepoint" issues by 2030, and high-level technological self-reliance by 2035, exporting Chinese research infrastructure and concepts globally. In this process, AI for Science will serve as the core enabler.

In the next five years, the world will see a wave of AI for Science infrastructure construction, with comprehensive iterative upgrading of data, computing, and experimental infrastructure. Traditional research literature databases and scientific software will be fully replaced by AI agents; scientific instruments will not disappear but will be comprehensively upgraded to AI-native intelligent devices. Industry consensus has formed: within five years, all R&D enterprises in biopharma, new materials, chemicals, and related sectors must fully transform into AI-driven enterprises, or their R&D efficiency will fall irretrievably behind.
Looking further ahead, before 2040, we expect to witness disruptive transformation in research: scientific discovery will become as simple as using a search engine, forming a standardized industrial pipeline paradigm; long-standing ultimate scientific challenges may finally yield to solution; the barrier to entry for research will drop dramatically, with high school-level ability sufficient to refine and define scientific problems; the population capable of participating in scientific discovery could expand from 245 million today to over 500 million by 2040; and the traditional logic of scientific value distribution will be reconstituted.
And China's goal of high-level technological self-reliance will be realized in the hands of our generation of research practitioners. Our core vision is to create AI scientists that surpass humans, solving the global shortage of research talent — to become the Anthropic of research.


