Please Help, AI! Solving the "Hard and Expensive to See a Doctor" Problem | A Conversation with Wang Guoxin: Chief Scientist, JD Health Exploration Research Institute
Ensure everyone has access to the same level of medical care.
Making equal-quality medical care accessible to everyone.

👦🏻 Podcast interview: Koji, Ronghui
🥷 Edited by: Starry
🧑🎨 Layout: NCon

Recently, several high-profile AI + healthcare companies in the United States have announced major milestones: OpenEvidence (medical knowledge search) has surpassed $10 million ARR, with tens of thousands of doctors paying to use it daily; Abridge (clinical documentation) closed a $250 million funding round; Tempus AI (oncology and precision medicine) listed on Nasdaq with a market cap that briefly exceeded $6 billion; and Hippocratic AI (healthcare-specific large language model) has reached a valuation of several billion dollars.
Together, these companies illustrate a clear trend: AI is rapidly reshaping the global healthcare industry. The same transformation is unfolding with equal intensity in China. The prominent Silicon Valley venture firm a16z has predicted that healthcare will be the industry that benefits most from AI.
This week, we invited Nico Wang Guoxin, Chief Scientist at JD Health's Discovery Research Institute, to share insights on the development and application of the "Jingyi Qianxun 2.0" large language model and the "AI Hospital." He discusses not only how, at the strategic level, AI products serve user health needs through JD Health's integrated medical ecosystem of examinations, diagnosis, and pharmaceuticals, but also analyzes the key directions and distinct emphases of leading American startups like OpenEvidence in the AI + healthcare space.
Healthcare is the most heavily regulated vertical, with the most sensitive data and decisions that literally mean life or death. The experiences and methodologies Nico shares today — how to identify real pain points, how to accumulate specialized data, how to compete for user mindshare at both the product and strategic levels — carry lessons for all vertical LLM industries, and we believe will offer valuable reference for anyone thinking through AI implementation.
Finally, Nico also shares his personal health management tips as a scientist — simple, practical, and accessible to everyone.

Listen on WeChat:
Listen on Xiaoyuzhou:


Lightning Round
👩🏻 Ronghui
Hello everyone, welcome to this episode of "Crossing." We've invited Nico Wang Guoxin, Chief Scientist at JD Health, to talk about medical large language models. Medical LLMs are a distinctive and representative case — we want to use them to explore whether all vertical AI ventures face similar challenges: Where does the data come from? How do you validate commercial viability? How do you actually execute? This is also the first time we've had a C-level chief scientist on the show, so I'll just call him Nico. Nico, please say hello to everyone.
👦🏻 Wang Guoxin
Hello everyone, very glad to be on "Crossing." Thank you.
👩🏻 Ronghui
Let's jump into the lightning round. Age?
👦🏻 Wang Guoxin
👩🏻 Ronghui
How many years have you been at JD Health's Discovery Research Institute?
👦🏻 Wang Guoxin
Three years.
👩🏻 Ronghui
What were you doing before this?
👦🏻 Wang Guoxin
Mainly search and multimodal technology.
👩🏻 Ronghui
Your MBTI and zodiac sign?
👦🏻 Wang Guoxin
ENFJ, Gemini.
👩🏻 Ronghui
Describe your current product in one sentence.
👦🏻 Wang Guoxin
We're building the Jingyi Qianxun medical large language model, and agent-based medical services built on top of it.
👩🏻 Ronghui
Can you share current revenue and profit figures?
👦🏻 Wang Guoxin
That's a killer question. Honestly, AI as a whole hasn't fully figured out its business model yet, but JD Health's revenue for the first half of the year exceeded 35 billion RMB, with net profit over 3.5 billion RMB. For us, this business is mainly about answering what the future of medical services looks like, and where the company is headed.
The Bittersweet Reality of Medical Data
👩🏻 Ronghui
As you mentioned, it's currently playing a forward-looking, communicative role. What's its function in the company's strategic planning?
👦🏻 Wang Guoxin
It's essentially core strategy. Healthcare is fundamentally a supply-constrained industry with extremely high service costs. Everyone has unlimited demand for health — I believe everyone wants to live longer. So the biggest question for AI here is whether it can expand supply. It's difficult, but not just JD Health — any company committed to health, and even at the national level, treats this technology with extreme importance.
👩🏻 Ronghui
Let's talk about the product you're building. You started in this direction roughly 10 years ago, began building internet hospitals in 2018 — would you say your current work is an extension of that?
👦🏻 Wang Guoxin
You could say that.
👩🏻 Ronghui
When you were preparing to build a medical-direction large model, what were your main considerations?
👦🏻 Wang Guoxin
As a company, the logic is to answer what the technology can deliver. JD Health initially developed certain AI capabilities because we handle over 500,000 daily medical service consultations — that's over 500,000 people seeking diagnosis and medication each day. Without AI triage and quality control technology, we couldn't achieve precise doctor-patient matching on one hand, and couldn't ensure online medical services remained compliant and legal on the other.
So JD's initial AI logic was to fit the business within a compliance framework while reducing costs. That was phase one.
In phase two, our team explored digital therapeutics and even some brain-computer interface work. Because healthcare isn't just online video or phone consultations — you also need to understand daily living conditions, test results. So we needed to extend data to both ends — health status before illness, recovery status after treatment. In this context, digital therapeutics technology became important. Then came the large model era.
The biggest appeal of large models is their human-like performance, especially instruction-following ability. In medical services, they don't just handle precise matching and compliance — they potentially achieve doctor-like service levels at dramatically lower costs. If capabilities advance further, it could even become a lifelong companion, like family, providing long-term presence.
This is the fundamental difference between this generation of large models and previous AI generations — it attempts to address healthcare's most basic pain point: supply capacity, achieved at low cost. In reality, some people enjoy expert-level service due to social status or wealth, but most cannot. Social stratification is even more severe in Western societies. Conversely, AI's value to the industry is whether it can expand supply at low cost. If achieved, everyone can access equal service levels, and everyone could potentially extend their lifespan by 3–5 years. This is the value medical AI can create for society and the country.
👦🏻 Koji
I've always liked this framing: large models may widen the productivity gap between people, but they definitely offer one universal value — emotional equity for all. In the past, many people suffered from psychological conditions without care, because mental health professionals were severely scarce. Large models can "see" everyone, hold space for everyone's emotions, even provide comfort. What Nico described about JD's health large model may be the generalization from mental health to overall health — bringing medical guidance and advice to everyone.
👦🏻 Wang Guoxin
At minimum, that's our vision, and we'll work hard toward it.

👩🏻 Ronghui
Medical models are vertical-domain models requiring massive amounts of specialized data. The advantage is that data should theoretically be highly standardized, but the difficulty is that collection and usage outcomes carry significant consequences.
👦🏻 Wang Guoxin
Healthcare has several characteristics — you could call it "bittersweet."
On one hand, years of national informatization initiatives and hospital performance evaluations have given hospitals relatively high digitalization levels — that's a fact. For example, medical records have strict standards, imaging has cloud data and quality control, and China's healthcare system invested enormous effort to cross the informatization threshold. That's the fortunate part of building medical models. If we were still at the stage of paper reports and manual registration, none of this would be possible.
On the other hand, there remain difficulties. First, medical data is often impossible to record completely, constrained by workflow needs, process requirements, or simply because there's no need to record data at massive completeness — a hospital's primary responsibility is still saving lives.
👦🏻 Koji
And what does get recorded is mostly distilled.
👦🏻 Wang Guoxin
Much of it is distilled data, and the quality doesn't necessarily meet our needs.
Second, model learning differs from human learning. Human doctors improve their cognition by inductively reasoning from cases, whereas models still rely primarily on raw data training, requiring large-scale inference or original data. But in medicine, such data simply doesn't exist. Often, doctors only record outcomes in medical records, not their thought processes. Humans acquire experience through extensive induction and oral transmission. This is why the concept of "self-learning" has emerged in recent years — when models learn to "learn" may be an important component of AGI.
Third, medical data is sensitive and complex. Case data is often scattered across different hospitals, and ownership is unclear: does the data belong to the hospital, the doctor, or the patient? There are also professional barriers. Take mutual recognition of lab test results — this was only promoted recently. Even CT scans can yield different results depending on the equipment or instrument, so consensus isn't automatic. This isn't necessarily hospitals seeking profit, but rather risk mitigation.
Therefore, medical models are highly complex, sensitive, and professionally insulated. Precisely because of this, they represent the most difficult and most valuable direction among vertical models. The difficulties are faced by everyone working in medical AI — we're all on the same track, and single-point technical advantages are unlikely to change the status quo.
An Internal Budget Request Formula: What Industries Merit Vertical Foundation Models?
👦🏻 Koji
Speaking of vertical models, I'd like to follow up. In your view, besides medical vertical models, what other industries need such specialized models? There's a prevailing view today that base models will become increasingly powerful and may eventually satisfy many generalized needs. But medicine seems to genuinely require large models. What other domains do you think need vertical foundation models? And can we abstract some underlying characteristics?
👦🏻 Wang Guoxin
Let me start with your last question — the abstract logic. Because I myself have to request budgets internally, I need to justify the necessity. I think the logic comes down to a few points:
First, is data in this industry relatively low-cost to obtain or simulate? If so, there's a case for necessity.
Second, is the business model sufficiently obvious? If it's too obvious, there may actually be no vertical opportunity.
We can map this along two dimensions: data and commercialization. For example, if data is obvious and low-cost to simulate, that means industry knowledge barriers aren't high, making it vulnerable to new technology substitution — a difficult position. Take education models: people say AI can make better math or English teachers. Today, it's generally accepted that AI's educational capabilities may surpass humans for language learning, because the knowledge is explicit and simulable, while also overcoming psychological barriers like nervousness about speaking foreign languages with strangers.
Another scenario is when knowledge isn't externally visible and requires enormous effort to govern, but the business model is sufficiently clear. This brings us to code. Current code models are often standalone models — one could even say Claude is a model optimized for code. In a sense, the so-called coder model is a vertical model. The reason general-purpose model companies also prioritize it is because the business model is too clear to ignore. What general-purpose company doesn't write code? In other words, our labor costs are too high — every company wishes it could replace the entire team with robots. That's a joke, of course, but it illustrates how externally obvious the business model is.
So to summarize, two conditions for vertical models:
- Whether the data possesses exclusivity, uniqueness, and urgency;
- Whether the business model is sufficiently clear and important to be impossible to abandon.
👦🏻 Koji
Right — I was reminded of our previous podcast guest, the founder of VAST, the 3D foundation model company. He mentioned that their training data is core competitive advantage. When asked about data sources, he said if he revealed that, he'd be exposing his most fundamental trade secrets.
Returning to medical foundation models: How do you view the changes AI has already brought to the familiar problems of "difficulty and high cost of seeing a doctor"? And what new changes will it bring in the next three to five years?
👦🏻 Wang Guoxin
I think AI has first changed information access equity. This is actually remarkable. Previously, when people got sick, their first instinct was to use search engines. But search engines' business model is auction-based ranking, which inherently promotes information mismatch rather than proper matching. We've discussed this with regulators as well.
Foundation models solve a critical problem: whether they can more faithfully respect physical reality. Today, in terms of business models, everyone is thinking, but few treat information matching as a business model. The pursuit is providing high-quality, trustworthy knowledge and information services. No one is challenging information matching anymore — it has shifted from "information matching" to "information absolutely correct generation." The best teams pursue exactly this goal.
So don't underestimate the significance of our shift from search engine models to Q&A and chatbot models. Behind it is the rapid improvement in information accuracy for ordinary people. From a medical perspective, "difficulty and high cost of seeing a doctor" presupposes health awareness. Everyone should minimize disease occurrence as much as possible. For example, gastroscopy penetration rates, physical examination rates, and lab test quality for people over 40 — these can all be further popularized through AI assistance, equivalent to educating patients and society as a whole.
👦🏻 Koji
So you mean that through conversations with AI, people will hear more health recommendations, like getting physical exams or gastroscopies?
👦🏻 Wang Guoxin
Yes, I think that's step one.
Step two is what's being solved today: whether models can solve triage, distinguishing user conditions. Can I identify whether you have a mild, severe, or emergency condition? Mild cases can receive standard solutions; severe and emergency cases can be quickly linked to medical resources. Previous AI systems couldn't do this. Before, you either made an appointment or went to an internet hospital to find a doctor. Now, if AI has been collecting your data all along, it can directly guide you to resources based on condition changes at critical moments — solving matching cost problems and reducing complexity and psychological barriers.
Third is the fundamental question: what level can AI-assisted diagnosis and treatment reach? This is the core battlefield for foundation models. If assisted diagnosis is sufficiently trustworthy, with humans only reviewing, then at least for common diseases, service capacity can extend to 7×24 hours, with gradually reduced human requirements — this partially solves the "difficulty of seeing a doctor" problem.
Finally, regarding "high cost of seeing a doctor" — this mainly concerns emergency and severe conditions. AI's greatest help isn't at the service level, but in therapy R&D. Massive medical papers are published daily — I can't even read them all myself, I can only look at the highest-level ones. Medicine is a lifelong learning profession, and doctors certainly can't keep up entirely. So we advocate for AI directed at doctors, continuously enhancing their capabilities. Improving doctors' skills is fundamental to improving overall healthcare quality.
Additionally, AI has become a core component in pharmaceutical and new therapy R&D. Unlike ChatGPT, it's not a toC product, but its impact runs deep. This year's booming innovative drug market is driven by Chinese companies' enhanced BD outbound capabilities, increased licensing deals, and faster R&D speed. This fundamentally pushes toward solving "high cost of seeing a doctor."
So to summarize:
- AI helps us improve health awareness and reduce misinformation.
- AI improves diagnosis for mild and moderate conditions, lowering psychological and service barriers.
- AI plays a core role in physician training, new therapies, and new drug R&D.
From a long-cycle perspective, these three points are AI's most important directions for changing "difficulty and high cost of seeing a doctor."

👩🏻 Ronghui
The vision you just described is beautiful. But beyond the shortage of physician resources, there's another barrier: for many people, using AI itself requires learning — they also need to be educated, which is equally challenging.
👦🏻 Wang Guoxin
I have a somewhat different view on this. A small case: once I was on a flight that was delayed, and an older gentleman next to me pulled out his phone to photograph the cabin. Curious, I glanced over and saw he was asking a chatbot: "What kind of plane is this? What's the model? Which seat is most comfortable?"
This made me realize that AI penetration is higher than many imagine. Especially in China. Although today's AI products aren't as visible on the C-side as in the mobile internet era, they've already demonstrated powerful capabilities in information services. Industry data shows this too — whether teenagers or people in their forties and fifties, AI usage rates are high, showing a "dual high" trend.
To some extent, AI chatbots are replacing search engines, directly providing answers rather than information. This is also why Google's stock was under pressure for a period — because AI can already directly provide knowledge and answers in users' eyes.
So I'm very optimistic about AI product penetration. I often imagine: if we went back three years to a world without AI, could we maintain today's lifestyle? I believe the answer is no.
👦🏻 Koji
I saw a similar question yesterday: If everyone else has AI and you don't, how much money would it take for you to live an AI-free life? I seriously considered it — even with 100 million, I'd have to think hard.
👦🏻 Wang Guoxin
Exactly — this is the cognitive gap created by productivity disparity. It can't be easily measured in monetary terms. I've lived through the transition from internet to mobile internet, that "unstoppable" trend — and the same is happening with AI now. Although AI's C-side performance hasn't fully iterated yet, penetration from B-side to C-side is already excellent enough. Otherwise we wouldn't see virtually every product defaulting to enlarged search boxes — behind that is precisely this transformation.
Jingyi Qianxun 2.0: Beyond Text, Three Core Evolutions
👩🏻 Ronghui
Back to vertical models. Your "Jingyi Qianxun" model went from 1.0 in 2023 to the latest 2.0 — could you introduce the main upgrades to our listeners?
👦🏻 Wang Guoxin
I think it shows in three aspects.
The first is changes in research methodology. At 1.0, we mainly used real knowledge data — papers, disciplinary articles, textbooks, and large volumes of real cases — which formed the data foundation. At 2.0, we've invested enormous effort in generating synthetic data.
So this time, Jingyi Qianxun 2 isn't just a model — we're also making our doctor-patient dialogue synthesis agent freely available to the industry. It's not open-source, but accessible through APIs. The industry's contribution is that through these APIs, people can simulate real doctor-patient dialogues as much as possible.
👦🏻 Koji
Does it directly use your simulated data, or can users also initiate their own simulations?
👦🏻 Wang Guoxin
It can initiate simulations. It's like a doctor — you can ask it anything, help it simulate consultations, and it can reproduce authentic doctor-patient dialogue from actual clinical settings, powered by our trained models. This is a new kind of understanding. Medical models often can't rely entirely on existing data, because that data is too difficult to obtain, so synthetic data or agent simulation is an inevitable path. The first change in 2.0 is the adoption of large-scale, high-quality synthetic data, which is made possible by JD Health's 490,000 daily consultations. We have the foundation to do this.
The second change is at the modality level. 2.0 supports imaging data, including CT, MRI, and X-ray. If you're limited to text modality in healthcare, you're far removed from the real world. Today, even if you've been coughing for over a week, a doctor will recommend screening — and for more complex diseases, imaging is a core diagnostic tool. So 2.0 represents a massive improvement at the modality level: it not only understands medical language, but also interprets imaging data with precision.
The third change is in reasoning. I've said before that I don't particularly like the word "reasoning," because in Chinese it carries ambiguity. Reasoning at the philosophical level involves human association and thought, while a model's "reasoning" is more like pattern learning — improving answer accuracy through compute power. It's not human reasoning.
In healthcare, the reasoning process must be verifiable. So we integrated with an evidence database. For example, if my reasoning leads to conclusions A, B, and C, I need to cite the source of evidence for each conclusion and grade that evidence — with top-tier journal articles or national guidelines at the highest level. Based on this, I make my diagnosis and judgment. So we call this "evidence-based reasoning," not just a way of thinking that consumes more tokens.
So synthetic data, multimodality, and evidence-based reasoning are the three major evolutions in 2.0, and why it deserves a new version number.
There's also a pretty cool demo showing that the reasoning process itself is multimodal. We can not only explain in text "because of A, B, and C," but also take an image, directly anchor to a lesion, and say "based on the state of this pulmonary lesion, I make the following inference." So its reasoning process involves multimodal interaction.
👦🏻 Koji
You mentioned that the first major upgrade was using large amounts of synthetic data — these being doctor-patient consultation dialogues. You said you use many methods to verify authenticity before using them for training. I'm curious how you verify this?
👦🏻 Wang Guoxin
This question can be answered comprehensively: all models in the medical field face challenges of data accuracy, model accuracy, and "how to verify." Our process works like this: during R&D, we build many evaluation datasets for comparison. But before any model goes live, it undergoes three stages of manual verification — this is costly.
The first step is in-house verification. JD Health has a large general practitioner team that evaluates from different specialty perspectives, measuring faithfulness, professional accuracy, fluency, consistency, and five or six other core metrics.
The second step is third-party verification. We collaborate with several major medical schools, which receive the model under partnership frameworks for secondary evaluation.
The third step is quality control committee verification. This committee consists of over 100 expert doctors from various regions who conduct independent assessments.
This process reminds me of the article OpenAI published, HealthBench. At the time, the CEO asked me about its significance, and I said it shows that even OpenAI needs doctors to verify medical models. HealthBench involved roughly 60+ doctors, including Chinese physicians, who manually wrote benchmarks and combined them with technical methods for verification. Our internal approach follows a similar three-layer model.
👦🏻 Koji
The synthetic data volume is enormous. With only 100+ experts, how do you verify so much data?
👦🏻 Wang Guoxin
The process can be understood as a funnel.
First, the funnel isn't filled in a day. Through continuous iteration, we can identify model issues and synthetic data bugs, making prioritization easier. Second, the upper layers of the funnel rely primarily on technical methods, trying to make machine evaluation as close to human evaluation as possible. The R&D team's goal is to minimize what flows to the lower layers while ensuring serious problems do reach them.
So you can think of it as a continuously iterating funnel. We don't verify every single data point. But from a probability perspective, large models are essentially Bayesian models — what we need to do is improve overall probability, leaving serious and error-prone cases for the lower layers while keeping straightforward ones at the upper levels, achieved through technical methods.
Where Can Vertical Models Outperform GPT on Specific Problems?
👦🏻 Koji
Actually, I have a big curiosity myself. Say I'm feeling unwell today — my first instinct is still to ask ChatGPT. Often I find its responses pretty accurate. So I want to know: as an 80+ person team spending so much time and effort training a medical large model, where can you do better than foundation models? Can you give a concrete example? If I asked our model instead of ChatGPT, would I get a more accurate, more comprehensive response?
👦🏻 Wang Guoxin
There are actually quite a few examples. I can approach this from two angles: single-modality and multimodal.
Let's start with single-modality. A true medical large model needs "expert-simulating capability" — it should think more like a doctor, not like an encyclopedia that covers everything. Patients may want to ask many questions, but from a doctor's perspective, what's more important is making quick judgments through a few key questions. General-purpose models typically draw on textbook knowledge, listing all possibilities and then exhaustively following up. But medical models should act like doctors, making rapid judgments based on core questions for specific conditions, rather than giving a long list of possibilities. It's not that general models can't do this — it's that it doesn't align with medical practice and ethics.
Now for multimodal. Take imaging. Many people use large models to translate articles, read papers — they work great. But if you ask a general-purpose model to interpret medical images, efficiency drops dramatically. Our model is specifically optimized for this: positioning, organ symmetry, sensitivity to small lesion detection. General models don't optimize for this kind of data because it's not their core business model, and there are data barriers. So in multimodal performance, the difference between us and general models is quite significant.
👦🏻 Koji
Multimodal I completely understand. But in single-modality, for common minor ailments like a cold, wouldn't the foundation model and vertical model give pretty similar responses? At what level of complexity or specialization do the differences become more obvious?
👦🏻 Wang Guoxin
Actually, fever is a great example. You can ask both the general model and our agent, then have real doctors review the answers — you'll see the difference. The general model will exhaustively list many possibilities because it's learned that fever is an extremely common symptom. But this isn't how medical practice actually works. A specialized model will conform more to doctors' habits and medical standards.
👦🏻 Koji
We can go back and ask this exact question to both the foundation model and our model, then put the comparison in the podcast shownotes for interested listeners to see for themselves.
(Editor's note: Regarding the comparison between "JD Health's large medical model" and "foundation models like ChatGPT" on the same medical health question mentioned in the podcast, the guest believes that 1-2 rounds of Q&A can't demonstrate the distinctive features. For those interested, we recommend searching for "AI Doctor" on JD to experience it yourself. We welcome your feedback after trying it out.)
👩🏻 Ronghui
I'm quite curious — OpenAI also does evaluations for healthcare, like benchmark scores. I saw you published Medbench results too. For ordinary users, the most intuitive thing might be seeing who scores higher. So how do you let users more directly perceive accuracy?
👦🏻 Wang Guoxin
This is a matter of experience, not just benchmarks. Frankly, benchmarks are more technical indicators that help us know what it takes to reach a certain level. But benchmarks and actual experience don't correspond 100% — this is also the difficulty in evaluating large models: everyone looks decent on paper, but there are still differences in real-world use. This involves the model itself, product design, even interaction design.
From our perspective, good experience means simulating expert service capability as closely as possible. But in medicine, what's most important is still diagnostic accuracy and effective treatment — this matters even more than experience. Of course we've also trained empathy capabilities, like teaching the model to express care and say considerate things. But this capability is general and can be separated from the medical model. The core of a medical model will always be the accuracy of diagnosis and treatment.
As for benchmarks, our internal attitude is: we can run them, or not. Often benchmark results don't align with our internal senior expert evaluations. Personally, I still trust real expert assessment more.
👦🏻 Koji
After all, benchmark evaluation dimensions are defined by another group of experts — it's just that their standards don't fully match doctors' standards.
👦🏻 Wang Guoxin
Right, and those standards are fixed.
👦🏻 Koji
Back to user experience. Yesterday at the JDD Conference (JD Discovery), I visited the Jingyi Qianxun booth and chatted with the product manager. I asked a similar question: in medical Q&A, what's the difference between foundation models and yours?
He gave a very interesting answer: in the JD Health app, user patient profiles are established, recording medical history and chronic conditions. So the same question gets different responses for different people, because it incorporates personal health information. Meanwhile, the app can also create family profiles — asking questions on behalf of children, parents. This seems like a small feature, but I think foundation models would struggle to do this unless extremely segmented. In vertical health products, this is quite valuable.
👦🏻 Wang Guoxin
Yes, I agree with your view.
👦🏻 Koji
Earlier you mentioned that emotional intelligence isn't one of the "holy grails" of medical large models. But for example, when Chuan Wang talked about "Baichuan building doctors," he emphasized the importance of communication: doctors must not only understand medicine, but also know how to communicate with patients and families. From your perspective, are you also trying to make AI more like experts in comforting patients and helping them rationally accept treatment plans?
👦🏻 Wang Guoxin
Internally, our evaluation system has two tracks: the experience track and the professional track. Treatment accuracy, consultation accuracy, and plan accuracy all fall under the professional track; communication skills, comfort, and communication ability fall under the experience track. Communication skills and professionalism aren't in conflict — often model capabilities can be orthogonal.
From an R&D perspective, we can use some data and algorithms to improve professional capability, and other data and algorithms to improve empathy, training them in one model and then triggering through prompts. Once a large model reaches a certain parameter scale, it gains generalization capability — unlike before when it had to memorize complete datasets.
So I agree that "communication is extremely important." But medicine is a high-trust domain where professionalism cannot be compromised. Communication is more like the infotainment system, while professionalism is more like autonomous driving — they have different logic and stability requirements. A model can already answer knowledge questions convincingly, but becoming a high-level conversational partner is difficult. In other words, the difficulty of being an internist is lower than being a psychiatrist, and the difficulty of being a psychiatrist is far higher than being an internist.
Improving empathy is necessary, but the hard part is figuring out how to evaluate and measure it. There's a saying in our industry: "When a metric can be measured, it can be optimized." Today, many models can mimic voices, and I believe voice mimicry is relatively easy. But if we made a digital avatar of Ronghui, Koji might find it convincing for the first few minutes, and then it would start feeling off. True human-like presence, high-level communication — that probably requires much greater investment and new technological breakthroughs.
So for me, this is a resource allocation question. Professionalism cannot be compromised, while we do our best to improve service quality. But I acknowledge that service quality still has technical hurdles.
👦🏻 Koji
Are you working on large models for mental health?
👦🏻 Wang Guoxin
We've considered mental health large models and have collaborated with top-tier mental health hospitals in China. This was a project supported by the Beijing Municipal Science & Technology Commission. The core was a mental health digital avatar — both the front-end avatar and the back-end model were developed by us, mainly aimed at alleviating patient anxiety and depression. Clinical trials aren't finished yet, but results so far have been positive.
However, from the model perspective, we didn't overly emphasize it as a mental health model. We're still primarily focused on common and serious diseases, with relatively less investment in mental health.
👩🏻 Ronghui
You've mentioned data acquisition several times. At yesterday's event, you also said you collaborate with many hospitals. How do you obtain training data?
👦🏻 Wang Guoxin
Our data comes from several sources.
First, partnerships with data centers. Medical data involves rights confirmation and compliance, and must be strongly de-identified and anonymized. Through data center partnerships, we obtain highly anonymized, compliant data. We recently signed with a national-level data center, with collaboration centered on large-scale multimodal models.
Second, our R&D approach is: internet data, JD Health proprietary data, and synthetic data form the baseline, then data center partnerships create a private data baseline. I believe we can obtain provincial-level data units through data centers, with massive coverage of the vast majority of common diseases.
Third, collaboration with individual top-tier specialty hospitals. They have long-term cohort data, mostly on difficult and complex cases. We currently collaborate with over a dozen top hospitals. For large models, our hypothesis is: combine data center training, then use small amounts of single-site data to enhance model capability. On compliance, data is obtained through scientific research cooperation agreements, with triple-party de-identification.
I believe most companies in the medical field will follow this path in the future.
👦🏻 Koji
You mentioned collaborations with hospitals earlier. What's the current attitude and assessment from hospitals toward our medical large models? Do they have concerns or reservations? Or are they generally supportive? Has any doctor or hospital director given you particularly memorable feedback?
👦🏻 Wang Guoxin
Actually, it's quite the opposite — they're broadly very supportive. I've been visiting hospitals for the past three years, and my sense is that support has been growing. Initially, it was some academicians pushing from a national perspective, then hospital directors, and now many department heads are trending in this direction too.
Collaborating hospitals have several core motivations.
The first is disciplinary construction. As national-level medical centers, they have a responsibility to develop their disciplines, and AI's ability to codify expertise and support disciplinary construction is inevitable — as is physician training — so they must participate.
The second is serving patients. They strongly want to extend their service capabilities and further transmit their experience — this is both an aspiration and a responsibility, so they're very willing to collaborate.
The third is that AI has already entered doctors' daily routines. Especially younger department heads — their understanding of AI often runs deeper than ours. I know one very young department head, a student of an academician, whose evaluation and awareness of different models' capabilities genuinely surprised me. The next generation of outstanding physicians growing up now will definitely use AI tools extensively to improve efficiency.
Of course, there's huge variation within the physician community, with completely different views on AI. My intuitive sense is: before last year, the emphasis was "can't make mistakes," but this year it's shifted to "mistakes are allowed, but they must be controllable and collaborative." They're more focused on which parts can be replaced, which can't, how to implement scenarios — they even proactively look for scenarios and solutions with us.
This is both exciting and pressure-inducing, because clinical scenarios are extremely variable, placing higher demands on model generalization.
👩🏻 Ronghui
In their feedback, which areas do they most hope AI can help with as soon as possible?
👦🏻 Wang Guoxin
Mainly three areas.
The first is patient services. Many doctors finish after seeing the patient, but medication tracking and pre-visit management still need support. Hospitals really want AI service bots that can accompany patients at low cost over long cycles, thereby improving cure rates or recovery levels. Diagnosis is just one decision — true health is in the individual's hands — so hospitals have massive demand for long-cycle services and transformation.
The second is department-level research. Research caliber and personnel development are very important to hospitals. Medical schools will definitely need to think about how to use AI to reduce learning costs and error rates. Many research-oriented hospitals want to build research platforms with us, entrusting their cohorts to us for automated mining, discovering new opportunities from patients, exploring new therapies.
The third is efficiency. Hospitals can no longer solve problems by adding staff — the cost pressure is too great. So they need "assistant" or "aide" type tools even more. Some hospitals have even proposed "doctor avatars," using digital humans to combine patient services with efficiency.
At the foundation level, demand is most concentrated in these three scenario types.
AI Hospital: A Calculated Move to Capture the "Future Health First Entry Point"
👦🏻 Koji
This time you also released another product — AI Hospital 1.0. Could you introduce what kind of product this is? For ordinary users, what help and value can it provide?
👦🏻 Wang Guoxin
The logic behind it is actually quite simple. We call it AI Hospital, and the core idea is: medical services have very strong professional attributes. In the past, we've developed many agents — psychologists, internists, pharmacists, nutritionists, and so on. Each agent can be optimized to the extreme at its specific point, which is the advantage of vertical agents.
The problem is: with so many agents, do we have users seek them out individually, or do we integrate them into a unified entry point? The latter is what we want to build. Whenever a user feels slightly unwell, we want them to think of coming here — that's the "mental entry point" we want to establish. So naming it AI Hospital is, in a sense, JD Health's exploration and contest for the future health entry point.
👦🏻 Koji
The future health entry point.
👦🏻 Wang Guoxin
Right, we could even call it "first entry point" — a stronger term, haha.
👩🏻 Ronghui
I feel like this product might include two directions:
First, in first-tier cities, user awareness is shifting from "seeing a doctor" to "health management." Many people proactively build health profiles, moving from passive treatment to actively reducing the possibility of getting sick.
Second, in non-first-tier cities, the medical resource gap is larger, and AI Hospital has the opportunity to become an entry point for accessing higher-quality medical services.
👦🏻 Wang Guoxin
Fully agree. The essential difference between AI and mobile internet is: mobile internet changed how people interact with information, while AI is more like B-side productivity. Although talking about B-side isn't exactly sexy in China, AI's core is indeed B-side empowerment.
If we project into the future:
- Large hospitals radiate to local areas through medical consortiums or mergers, handling complex diagnosis, treatment, and rehabilitation services.
- At the more granular community level, AI assists local doctors, responsible for screening, triage, consultation, and referral.
- AI can also extend services at low cost, connecting effective medical resources.
Looking at China's future, as population aging and regional disparities grow, this model will very likely take shape. Of course, payment models will change accordingly, but that's another topic.
👩🏻 Ronghui
So how do you plan to actually implement this? Especially getting it to the people who need it most, rather than leaving it at the product level?
👦🏻 Wang Guoxin
Actually, JD Internet Hospital itself has been doing this. The underlying logic of internet healthcare is cross-regional medical resource matching and 7×24 availability. AI isn't an entirely new story — it's a further enhancement built on top of existing internet healthcare. In other words, AI healthcare is the natural extension of internet healthcare.

👦🏻 Koji
Speaking of "Jingyi Qianxun," it's open source, right? Could you specifically introduce which parts are open source, and why you chose to open source?
👦🏻 Wang Guoxin
Let me start with "why." Healthcare is a trust-driven industry. Through open sourcing, we can bring ecosystem partners in, demonstrate our technical capabilities, let outsiders try the model and give feedback, thereby feeding back into R&D and ecosystem building. This is something we must do.
Our open source commitment is also quite substantial: not just the model, but also training code and partial training data. We want participants to truly reproduce our work, not just receive a result.
👦🏻 Koji
So the core goal of open sourcing is building trust. After open sourcing, do you feel this goal has been achieved? Have you received feedback from the community or partners?
👦🏻 Wang Guoxin
Main feedback comes from research institutions, including universities and hospitals. Especially for smaller-scale models, many partner hospitals proactively test them. This has greatly helped us push forward specialty collaborations. Open sourcing lets others see that we're a team that genuinely does things, which strengthens trust.
So the biggest gain is: hospitals and research institutions are more willing to collaborate with us. Trust itself is priceless, and open sourcing played an important role in this process.
👩🏻 Ronghui
You mentioned earlier hoping the product can occupy user mental space. I do think it's possible. Especially if more and more users get accustomed to asking ChatGPT or other chatbots medical questions, that would impact your entry point advantage, and could even disrupt the entire business model. So what kind of synergy do you expect between AI models, AI Hospital, and your existing business model?
👦🏻 Wang Guoxin
When discussing AI business models, I think a few things are most valuable. First is "high-reliability replacement." Even in just a very narrow domain, if AI can achieve 99.9% reliable replacement, that's extremely important. Second is "connection." Whether AI can become a better bond, connecting consumers with services.
In healthcare, both opportunities exist. Combined with JD Health's model, we must return to the group's core logic: we are a supply chain-driven company. That is, our advantage lies in providing the highest quality products and services at the lowest cost. AI can play a huge connecting role in this. So for us, entry-point products must be contested and pushed forward.
JD Health isn't just an internet company — we also have physical medical institutions and home-delivery service capabilities. In many cities, for example, we can deliver medicine to your door within 30 minutes. We have health examination centers, hospitals, and pharmaceutical supply chains. In this process, AI's role is to connect these service capabilities and provide patients with a complete set of personalized solutions.
So this isn't a question of "whether to do it," but "how to do it." Future competition will definitely shift from single-point chatbot battles toward the combination of "chatbot experience + backend service capability," ultimately coming down to whether you can deliver user satisfaction. The core of healthcare is efficacy — only with efficacy can you survive.
👩🏻 Ronghui
On the backend service side, would you worry about it affecting the frontend chatbot's information delivery?
👦🏻 Nico Wang
No. We make our backend services as atomic as possible. For example: a nurse visiting your home for an examination is one atomic service. The model's job is: based on the patient's current condition and the conversation, determine whether to trigger this service, how much it costs, and whether the patient is willing.
The model solves the information side; the backend supply chain handles execution. Our supply chain isn't just goods — it includes services. These capabilities are the foundation of JD Health's user mindshare. Without them, we'd just be an internet company floating in the air.
JD's core mindshare is "high-efficiency, low-cost service capability." On this foundation, we have the chance to build entry-point products. Though entry products are difficult, all AI companies are thinking about this right now.
👩🏻 Ronghui
What about other foundation model companies? Could they extend products or services based on users' medical consultations in their chatbots?
👦🏻 Nico Wang
Many foundation model companies are very focused on the health track. Especially some large chatbots — a significant portion of their traffic is health-related. This is very similar to the logic of how search engines captured mindshare back then. Many people still use large models as search. So for them, this is a track they really want to pursue. But the key questions are: does this industry actually have moats? Is commercial maturity sufficient?
For JD, we see them more as partners than competitors. Also, JD released its own general-purpose chatbot at the JDD conference — the newly upgraded JoyAgent 3.0. I certainly hope it can quickly establish itself in the market and drive internal industrial synergy within the group.
👩🏻 Ronghui
Nico, could you tell us about AI + healthcare across broader markets — the United States, Europe, and China? In these markets, what innovations or success stories are worth noting? For example, OpenEvidence in the U.S. has done quite well in both fundraising and revenue.
👦🏻 Nico Wang
Healthcare AI doesn't transfer across overseas and domestic markets as easily as other industries. The key difference lies in payer logic and healthcare systems. China emphasizes efficiency and equity. Though people complain about "difficulty and high cost of seeing a doctor," the problem would be even more severe in the United States.
OpenEvidence can commercialize in the U.S. largely because doctors have high incomes and strong demand — they're willing to pay subscription fees for tools. But in China, our model is free. For example, the evidence library we just opened has no subscription fee at all. This is the "orange that grows in Huainan is an orange, but in Hubei it becomes a trifoliate orange" situation.
That said, some overseas models are worth watching. For instance, in the U.S. there are quite a few AI-driven telemedicine + specialty pharmaceutical service companies. Take Hims as an example — it positions itself as "making people more beautiful, making them better." Essentially it relies on specialty pharmaceutical supply chains, but front-end customer acquisition and services are all AI-driven, continuously giving users health recommendations.
Overall, medical solutions still tend to be pharmaceuticals, devices, or lifestyle changes. AI can help hospitals improve services, and it can also help pharmaceutical companies, device manufacturers, or digital therapeutics companies serve users.
Beyond that, there's the AI service model for hospital clients. Domestic hospitals already have relatively high IT penetration, but procurement cycles are very long. We also have a smart healthcare division, such as JOY DOC, aiming to use AI to transform medical and patient services.
Finally, there's ToG — facing government, serving medical insurance and health commissions. This is relatively rare in the United States.
To summarize, commercial opportunities in the China market will ultimately return to the patient services track — it's more suited to local soil.
Advice for Ordinary People: How to Use AI to Live Better?
👦🏻 Koji
Nico also mentioned that these past three years, you've been frequently visiting hospitals, immersively thinking about how AI plus healthcare can truly help everyone become healthier. So if we return to a friends-chatting scenario — we meet today, and I ask you: you've researched so much AI and healthcare, can you give us some advice now? The kind that's small, highly actionable, that listeners can take and use right away to make themselves healthier. How would you answer?
👦🏻 Nico Wang
From a long-cycle health perspective, there are mainly two influencing factors: first, personal performance on chronic diseases and immunity; second, risk of serious illness. As age increases, investment in health checks and early prevention is definitely worthwhile.
For example, after 35, I believe you should set aside a fixed budget every year for personal and family health. It doesn't need to be a lot, but it must be fixed — consciously using economic means to push yourself into action. Statistically speaking, this actually saves money, because many diseases are curable when caught early. The key is to set a budget, and within that budget find the best medical services.
👦🏻 Koji
Right, I find this very interesting. Take the money out first, then figure out how to spend it — forcing yourself to take action. It's more effective than simply saying "everyone should get checkups," because you have to spend the money, otherwise you have to punish yourself at year-end.
👦🏻 Nico Wang
Exactly. Of course not everyone needs a gastroscopy or colonoscopy — that depends on family history and personal risk. I'm just giving an example. The core logic is: set the budget first, then make health investments suitable for yourself. Health itself is quite anti-human-nature — people often only realize its importance when they've lost it.
👩🏻 Ronghui
What about people under 35?
👦🏻 Nico Wang
The same principle applies. Especially pay attention to family history and your own condition. Some things can be persisted with long-term, like monitoring blood pressure and blood sugar. This seems simple, but is very helpful for early detection and early intervention. Take diabetes — the outcome difference between early detection/early control versus late detection/late control is enormous. Many diseases have solutions in early stages; once you miss the window, you can only mitigate rather than cure.
Investor Perspective: How to Judge a Vertical Large Model Company?
👦🏻 Koji
We've talked a lot about healthcare large models, but many listeners aren't in healthcare — they work on various vertical models. How do you think healthcare large model experience can transfer to finance, law, and other fields?
👦🏻 Nico Wang
Actually the technical approaches in these industries are quite similar. Healthcare, education, law, finance — fundamentally they're all about modeling and optimizing in complex situations. For example, personalized learning paths in education, multi-step reasoning in law, investment portfolio recommendations in finance — all require handling highly structured data and performing fused reasoning. This is why healthcare experience transfers easily to these industries.
👩🏻 Ronghui
If you were an investor, how would you judge whether a vertical large model company can succeed? What metrics do you care about most?
👦🏻 Nico Wang
I mainly look at three things:
- Industry knowledge depth: Do they have real data moats and professional knowledge accumulation? This is a 0-or-1 question — without moats, it doesn't work.
- Size of commercial opportunity: Can't be too large, or big companies will enter and you'll have no chance; but also needs near-term monetization potential that makes sense.
- Future commercialization tone: Will they go API payment, product payment, or sales-driven in the future? This depends on the business partner's capabilities and thinking on the team.
👩🏻 Ronghui
Thank you so much to Nico for coming on Crossing today and sharing so much experience with healthcare large models. This field has both value and intense attention and expectations. We also hope AI can truly let more people enjoy the medical fruits that technology brings.
👦🏻 Koji
Thank you.
👦🏻 Nico Wang
Thank you both, bye-bye.
