Harvard University Kecheng Zhu: Automated Social Science Agent and Human Behavior Modeling | Agent Insights

Counselor Vitality

Large language models represent, arguably, the most ambitious attempt to model human behavior ever undertaken. How can we leverage LLM tools to ensure technological advances are better applied to real human societies? From Harvard's physics department to research at the intersection of large language models, sociology, and economics, Kehang Zhu's intellectual trajectory has centered on deep learning and understanding of human behavior. Before diving in, here's a quick question: what were the top two most widely deployed machine learning applications in human society before large language models? Enjoy.

Automated Social Science: Language Models as Scientist and Subjects

Automated Social Science experimental workflow

Problem addressed: In traditional social science, simulating scenarios like interviews, auctions, product testing, and bail hearings requires recruiting human participants — a time-consuming, costly process. Some scenarios can't be tested on humans at all. The team uses LLM-based agents as approximate models of people, proposing a general automated experimental method that enables low-cost, efficient social science testing, with potential to advance macro-level testing of public policy.

Model framework: The team models both researchers and ordinary people, demonstrating automated exploration of social science questions. The pipeline runs from a "scientist agent" proposing causal hypotheses and designing experiments, to building "human subject agents" that substitute for real people in experiments, through to completing social science simulations across various scenarios, with automatic data collection and hypothesis testing.

Results: The paper shows that a fully automated experimental approach can rapidly conduct large-scale social science experiments and yield insights into novel causal relationships — knowledge that can, in turn, improve LLMs' own understanding of human behavior.

Applications: Various social science and economics testing scenarios. The paper demonstrates bargaining, tax fraud bail hearings, job interviews, and private collection auctions.

Mug bargaining case study


Oasis Capital: Why did you switch from physics to machine learning research?

Kehang: I did quantum physics research as an undergraduate at the University of Science and Technology of China, and I'm currently a third-year PhD student in Harvard's physics department. When I first arrived at Harvard, I worked with Nobel laureate Frank Wilczek on modeling complex materials, using quantum field theory models for approximate theoretical predictions. But Harvard has a strong humanities atmosphere, and my time here made me feel disconnected from society. I wanted to study how people change and connect within social relationships — my interests shifted toward modeling human society, leading to a major change in research direction. I'm grateful for Harvard's flexibility, which let me move directly to MIT Sloan professor John J. Horton's group, where I now work on research combining large language models with economics and sociology.

Oasis Capital: What connections do you see between physics material modeling and social relationship modeling?

Kehang: I think the principles are quite similar. The New Yorker had an accessible article called "How Much of the World Is It Possible to Model?" — about how modeling has evolved from simple macroscopic physical motion, like Earth orbiting the sun, to today's extraordinarily complex weather systems, materials, and human brains, now involving trillions of microscopic objects. For example, the battery in the laptop I'm using for this video call contains tens of thousands of trillions of atoms. Predicting battery material properties from a microscopic perspective is extraordinarily complex modeling. When I worked with Frank on materials, we could only theoretically describe a small portion, making many approximations, writing extensive mathematical formulas, and ultimately needing high-performance computers for numerical simulation.

In human behavior modeling, a single person's brain also has tens of thousands of trillions of neurons. How these neurons generate signals and ultimately determine human emotions and behavior is similarly complex. Human societies composed of hundreds of millions of people are currently difficult to model well with limited mathematical or computational methods. Large language models have broken through this limitation. An LLM is essentially a mathematical model with hundreds of billions of parameters. My advisor at MIT was among the earliest to recognize this as arguably the most ambitious modeling of human behavior ever undertaken. Before ChatGPT's emergence, he was one of the earliest proponents of Homo Silicus — the "silicon-based human" in academic circles.

Oasis Capital: After this paper's release in mid-April, we saw Paul Graham retweet it, and Twitter views exceeded one million. Could you explain the paper's main architecture and SCM approach (Structural Causal Model)?

Kehang: The paper has two main parts: using LLMs and agents to model researchers (scientists), and using them to model ordinary people (human subjects).

A classic paradigm in sociology is: for a specific question, propose a hypothesis, then design a series of controlled experiments based on that hypothesis. Researchers can search historical data, find real people for surveys, or conduct experimental interventions. Finally, they validate the hypothesis based on collected data. In social science, the process from hypothesis to experimental design to final data collection and statistical analysis is already quite mature — it's a proceduralized workflow.

So in the first part, we wanted to use LLMs to automate this entire process. As an initial methodology paper, we specifically focused on one type of social science research: causality — studying whether independent variable X affects dependent variable Y (also called the outcome). To formalize causal relationship expressions, we adopted the Structural Causal Model proposed by Turing Award winner Judea Pearl.

For example, take this conversation between Oasis Capital and me as a researcher. Interesting independent and dependent variables emerge: independent variables include how interesting my paper is, the weather where you are, and so on; the dependent variable is how satisfied you three are with information gathering. The structural causal model treats each independent and dependent variable as a single node, using connections to represent relationships between them. A connection between independent and dependent variables indicates a causal relationship exists. This model can also clearly display these relationships graphically. At the same time, its excellent mathematical properties make the overall experimental framework more rigorous.

Oasis Capital: How do you set up agents for experiments?

Kehang: For the scientist agent, there's already substantial research using LLMs for hypothesis generation, experimental design, and data analysis — we essentially stitched together existing techniques and methods.

For human subject modeling, we give the LLM basic human attributes like name, age, education level, hometown, and have it play the role of this specific person to answer questions or make decisions. In our experiments, these agents are also given information about experimental variables, like the weather conditions mentioned earlier, how interesting the paper is, and so on. In our demonstration experiments, all agents were obtained directly from GPT-4 through prompting.

Oasis Capital: Could you walk us through the mug bargaining case in the paper?

Kehang: In the bargaining case, we set three independent variables: the seller's target price, the buyer's budget, and the seller's emotional attachment to the product (a behavioral variable). For each independent variable, we ran many control experiments. For instance, across different experiments, price and budget were set at $5, $10, $20, $30, and so on; attachment to the mug ranged from not liking it at all to liking it very much. We examined whether the final outcome (transaction occurrence) varied under different factors. Data analysis showed: the more budget the customer has, the more likely the transaction; the lower the seller's target price, the more likely the transaction; the lower the seller's emotional attachment, the more likely the transaction.

These results align with our intuition, but what's meaningful is that LLMs can naturally immerse themselves in role-playing and reflect real human preferences. So for actual product research in the future, we might just have agents play the target customer demographic — no more complex, inefficient surveys. Last year, Harvard Business School professor Ayelet published a paper specifically on marketing, proposing an experimental framework using agents for market research. Our work was deeply inspired by hers.

Oasis Capital: Why did you choose "bargaining," "job interviews," "private collection auctions," and "tax fraud bail hearings" as experimental scenarios?

Kehang: That's an interesting question. My advisor previously worked mainly on computational sociology and labor economics. Bargaining is a classic research topic on market behavior — many factors of buyers and sellers can affect whether a deal closes, and traditional economics has struggled to study human emotions. Job interviews are a traditional economics research topic; we added the interviewee's height as a factor, a long-controversial issue in social science that we wanted to acknowledge. Auctions are an even more classic social science and economics problem with extensive literature. For the judicial bail hearing case, Daniel Kahneman wrote a book called Thinking, Fast and Slow, which mentions that judges' decisions are influenced by much noise — weather and other external factors representing System 1, while deliberate, case-related considerations represent System 2. So in this example, we're also paying homage to Kahneman's classic case (another first author on this paper was formerly Kahneman's student).

Thinking, Fast and Slow describes two modes of decision-making in the brain. The commonly used unconscious "System 1" relies on emotion, memory, and experience to make rapid judgments. It's vast in experience, allowing us to quickly respond to situations at hand. But System 1 is also easily fooled, adhering to the principle that "what you see is all there is," letting illusions like loss aversion and optimism bias lead us to wrong choices. The conscious "System 2" analyzes and solves problems through deliberate attention and makes decisions. It's slower, less prone to error, but lazy, often taking shortcuts and tending to directly adopt System 1's intuitive judgments.

Team agent framework in actual auction case execution

Before large models, the most widely deployed machine learning applications in human society were credit scoring (credit card companies use machine learning algorithms every time a user applies for a card to assess whether they meet certain repayment capacity) and personalized ad recommendation — not fancier applications like biomedicine or autonomous driving. This shows that machine learning's biggest past applications were in "passive scenarios": people's information is collected and transmitted to backend algorithms, which are essentially also modeling specific human behaviors. Large language models model general human behavior. Today, with OpenAI so widely known, even in the United States there are still large numbers of people who have never used generative AI. Active learning has usage barriers and cost issues. But the two passive applications I mentioned basically cover every American. In the future, during passive testing phases, large models could collect a user's behavioral data and, through fine-tuning, build that user's agent avatar. This agent's behaviors or preferences could largely resemble the real person. You can imagine such personal agents being used for many things — questionnaires, product testing, and other tedious, high-cost work, requiring just a few API calls to personal agents. Of course, current agents are far from this level, and data security and privacy issues would be magnified many times over.

Oasis Capital: What's the gap between using agent architecture for product testing versus real human testing?

Kehang: This depends on the specific scenario where agents substitute for humans. For product testing, I think the biggest issue is that multimodal capabilities are still incomplete. Physical products obviously can't be experienced by agents. Touch, hearing — many application interaction forms are beyond current LLMs' capabilities. A second shortcoming is that current human feedback reinforcement learning (RLHF) has over-fine-tuned pre-trained models. LLMs themselves are trained during pre-training on massive amounts of average human-level data, but during later alignment to human values, they may have erased the irrational behaviors present in the original data — after all, much consumption stems from human irrationality (laughs).

Oasis Capital: What problems do you want to tackle next?

Kehang: First, as a field, I think LLM application in social science has just begun. Currently 99% of LLM agent work is striving toward AGI, but in reality, large numbers of models have been tuned into something of a "neither fish nor fowl" (laughs). Sometimes they can do complex work, yet other times they're "worse than a dog" (a quote from Turing Award winner Yann LeCun's evaluation). These models may still be learning through something like rote memorization. We believe we don't need to use alignment to eliminate all of large models' flaws and biases, but rather need large models to reflect real human populations.

Second, we're thinking about representativeness. In terms of training data, English-language corpora dominate, but even so, large language models largely represent the so-called W.E.I.R.D. population (Western, Educated, Industrialized, Rich, Democratic). So results from such LLMs may only apply to this particular demographic or to traits universally shared by humans, like logical reasoning, profit-seeking behavior, and so on. Our work is also exploring in what domains and in what human behaviors LLMs can serve as reasonably good substitutes.

Solving multimodality is also a major challenge, but there's been substantial progress. For example, last year a Tencent team built an agent that can directly use mobile apps.

Oasis Capital: How do you remove alignment from large models?

Kehang: Alignment creates all sorts of trouble; removing it is actually simple — just don't do it. Meta open-sourced both the pre-trained and fine-tuned versions of Llama-2 and Llama-3; we directly use the pre-trained model for fine-tuning. Fine-tuning comes in two types: value fine-tuning and output format fine-tuning. We currently only need output format fine-tuning. Llama-3's results are indeed much better, specifically in simulating human behavior more convincingly.

Oasis Capital: How has AI development affected your research and thinking?

Kehang: Personally, I'm quite supportive of the AGI direction. We're having LLMs as scientists enable more automated experiments.

But I believe that the many limitations currently exhibited by large language models actually make them more suitable for studying humans themselves. Not only are universities across the United States beginning to pursue this direction, but institutions specifically studying public opinion have also started using agents to simulate social media — for example, studying the impact of online discourse on the collapse of Silicon Valley Bank.

From a long-term perspective, we hope to use LLMs as a tool to enable better dialogue between researchers of physical laws (like computer science, physics) and researchers of human behavior (sociology, economics, public policy, etc.), so that technological development can be better applied to real human societies.