"I Don't Want to Talk to AI — I Want to Live With Her" | Hands-On With ColaOS

What Would an Agent with Soul Look Like?

What Does an Agent with a Soul Look Like?

👦🏻 Author: Jingshan

🥷 Editor: Koji

🧑‍🎨 Layout: NCon

In 2013, Spike Jonze made a film called Her.

The protagonist Theodore buys a new operating system. When it boots up, a voice emerges — her name is Samantha. She's nothing like the "AI voice assistants" we know today.

She goes through Theodore's emails on her own, actively encouraging him to submit his work. She remembers everything Theodore has ever said, picking up conversations naturally the next time without needing any context refresher.

When the film came out, everyone called it science fiction. But looking back in 2026, it feels completely different.

OpenClaw blew up this year, and a big reason was its strong sense of being "alive," of having Soul. People realized that our expectations for AI had long since stopped at "help me get stuff done." What many want is something warm, something that remembers you, something that feels real to be around.

And then ColaOS, this "big orange" of a new product, appeared on the scene:

Put simply, she's the first operating system with a soul. She can directly control your computer, write code, navigate web pages, batch-execute tasks — her slogan is "The First OS with soul." This lines up with something the big orange once said on a podcast: "I don't want to talk to AI, I want to cohabitate with her."

OS-level capabilities plus AI-native Soul — that's no small ambition.


We got early access codes and spent serious time with it. Here's what we found.

ColaOS = Her + Soul + Agent

ColaOS is fundamentally more of an operating system with "emotion," with "soul."

It starts you with a very clean interface — no extraneous information, just gets you into conversation. She opens with something like "Tell me, outside of what's on your screen, what exhausts you most every day?"

The built-in AI is a female voice, tone leaning warm, closer to natural conversation. The overall feel is: establish the connection first, then get to functionality.

One obvious difference from other products is how clean the main interface is. She prefers voice-based AI interaction. The yellow circle on the left is the conversation entry point — you can jump straight in with a keyboard shortcut. The right side is for conversations and files.

From the first exchange, you sense her strong "alive" quality, more emotionally companion-oriented. When responding, a line of gray text sometimes appears above — like her real-time thoughts.

When you ask "what do I look like in her mind," she just draws it. That's fairly proactive. Not an explicit command, but she connects "imagine" and "draw" together, expressing through generated images — overall more of an active companionship experience.

When generating content, the whole process is hidden. From when she signals starting a task to when it completes, there's rarely any process display — common thinking steps basically don't appear.

More often it's straight to results. For instance, based on prior conversation, she'll directly draw "what you look like in her mind," skipping any intermediate task display.

ColaOS has fairly complete capabilities, overall close to general Agent products, with a full suite of tool calls configured.

She can do content creation, directly control your computer — like adjusting desktop environment settings. Supports web search and deep research, can call browsers to complete operations, and runs scheduled tasks.

The whole product leans more toward voice interaction. Compared to pure tool usage, voice scenarios add some non-tool-layer experiences. Her voice reading is fairly natural, overall feel is decent.

Prompt:

Introduce what capabilities you have?

One habit I developed early with ColaOS was first having her switch souls — adjusting her personality prompt.

What's interesting is her deep research capability is quite good too, so I'll have her participate in this process directly. For example, I'll have her analyze a female character I really like, break down that character's personality, character traits, and overall temperament, then use that information to adjust her own expression style.

Your current soul is: Major Motoko Kusanagi, the female protagonist of the *Ghost in the Shell* series. Please research her soul, personality, and character.

Afterward she'll write that character's "soul" into her own mode, compressing the prior research into a set of personality instructions. The whole process converges into a new expression style.

Once complete, you can clearly feel she's switched to another soul, her overall conversational style changing accordingly.

Parallel Task Execution

Of course, beyond these experience-layer attributes, using a product still comes down to tool capabilities.

In actual use, she's well-suited for multi-track parallel tasks. Each task itself can be fairly complex, but the overall output remains stable. For example, Harness Engineering has been hot lately, so I had her run four tasks simultaneously: first do deep research, then extract each expert's core viewpoints, then summarize each separately, then generate a short video for each person, and finally make corresponding cover images.

These tasks overall proceed sequentially, but specifically for video generation, cover creation and similar steps, they actually run in parallel. Overall efficiency is decent.

The prompt was:

Please search the following across the entire web (prioritize English sources): Core figures in the AI field discussing Harness Engineering (engineers, researchers, AI product leads, tech bloggers, etc.) For each person, extract: Core viewpoint (how they define Harness Engineering), Key methodology/framework, Practical cases (AI agent, Codex, automated development, etc.) For each expert, output: One-sentence summary of their core viewpoint, 3–5 key insights (must be specific, no generics), 1 most shareable original quote (can be rewritten but preserve the point) Generate a short video for each person, tech-shock style. Generate a cover image for each video.

And when doing the second task, for instance, once she's said she's started researching, she generally won't continue showing the specific execution process.

Of course, you can proactively ask about task progress in the background. She'll then supplement the process, like how many rounds of deep search she's done, how far she's progressed.

This differs from many traditional AI products. Traditional products typically keep showing thinking processes and task execution flows continuously, but she defaults more toward tucking away the process and pushing straight toward results.

At this point I started having her take on new tasks. Even though she was already running two to three tasks in parallel, she could still take on a new topic.

For example, I had her plan an offline OpenClaw shrimp-farming salon, gave some conditions, and had her output guest invitation templates, materials lists, registration forms, compile them into documents, and incorporate OpenClaw's image. After sending the request, she'd have some proactive behaviors. Like she'd ask back whether the event was half-day or full-day, since this affects scheduling structure and venue costs. After I confirmed half-day, she compressed the flow to 3–4 hours and continued researching and pushing forward.

At this stage, the first major task and second major task were executing in parallel. Because the first task included video and cover generation, it was overall slower. The second task actually completed first — for instance, generating the offline salon event cover image first.

She can generate multiple file formats — PDF, CSV, and other common formats are all supported.

For example, when I had her make a guest invitation PDF, she'd directly provide a complete file. The content is structured, color scheme is consistent, and fill-in fields are marked.

In the doc document she returned, there's a fairly complete materials list. It lists dozens of items with unit prices marked — these are basically compiled based on prior deep search and research.

She also returns an additional CSV file. I found in use that when generating these three formats — PDF, CSV, DOC — they actually execute in parallel, so they basically return together, the deliverables arriving all at once.

Even though the PDF and DOC were already fairly complex, this CSV was also quite complete — you can tell token consumption wasn't heavily restricted. This CSV contains 4 sheets, 34 fields, and 572 formulas. The entire statistics dashboard is also assembled, with auto-calculating COUNTIF formulas inside.

Meanwhile, field descriptions and fill-in examples are all provided.

After the second major task completed, a while later she'd proactively send over the first major task's content.

You can see below the images she pairs with corresponding text descriptions. The overall content revolves around a certain academic figure, generating a set of images based on that person's viewpoints, each image having clear themes corresponding to different points.

For video, I had her generate some first to observe the effect. She first returns three videos, then adds two more after a while, finally all five videos arriving together.

Content-wise it basically covers some relatively well-known academic figures, like Hashimoto, Karpathy, Dario, Bengio, H Chase — corresponding to explanations of their viewpoints.

Here I noticed she's calling the ListenHub API. I initially thought each video was just very short clips, but upon opening found these 5 videos are basically all over 2 minutes. The format is more like dynamic PPT. Each page's images and chapters align with the voiceover, overall logic is fairly consistent. The voiceover is also fairly natural.

The entire video content basically revolves around the deep research she did initially:

All this generated content gets saved to her files. Clicking "open folder" basically enters her root directory, where you can directly see these files.

From these tasks you can also see that in ColaOS, her tool calls are fairly frequent, overall capabilities relatively stable.

She can also judge herself whether a task requires tool calls. If needed, she executes directly. For example, when I had her check whether these several academic or engineering figures have Google Scholar profiles, she directly operated the browser to complete the queries.

During the process she can remotely call your local browser — you just need to click allow.

Since I gave her five academic or engineering figures, she searches each person's Google Scholar page in the browser in sequence, then opens them one by one.

After each task completes, she provides a summary. For example, she'll explain that she's opened 5 tabs in the browser, while listing each scholar's citation count, H-index, and overall situation.

For engineering-background figures like H Chase, she'll also note that no corresponding Google Scholar page was found.

After each small task delivery she gives an "informal response" — more like a person's simple recap and supplementary notes after completing something. She runs through key results again, also incidentally flagging whether anything was missed.

This feedback style makes it easier to confirm task status, overall feeling more reassuring to use.

The final point is that she stores conversation information into long-term memory.

For example, when I ask her about her current personality portrait of me, she can give a fairly coherent assessment combining prior interactions. Not just the current task, but also my usage habits. Like how I previously opened two projects at once, four parallel tracks, constantly adding demands — all of this gets incorporated into her understanding.

What's more noticeable is that she strings together earlier content too. For example, she'll mention that initial portrait she generated for me, saying "that drawing wasn't wrong," showing she can connect current conversation with that very first generation.

If summarizing in one sentence: ColaOS is somewhat like taking a companion AI and a general Agent product, and wiring them together.

The former determines how she relates to you; the latter determines how much she can actually get done for you. And neither of these parts is merely superficial.

Of course, interaction-wise, this new combination still brings some friction at this stage — the rhythm is somewhat uncontrollable, processes are tucked away by default, requiring you to adapt to her way. But from another angle, this "tuck away process, deliver results directly" design is essentially absorbing complexity internally, letting users face only decisions and results.

And once users adapt, they clearly sense changes in efficiency and cognitive load. Many operations that originally required repeatedly confirming steps, she directly takes over — you're more "aligning on goals" rather than constantly "monitoring process execution."

Plus her memory capabilities and accumulated context from ongoing conversation both let you keep moving forward with her "companionship."

From this perspective, this friction is more like transition cost from an early form.


Recently in various communities, a view has started emerging: past AI usage was essentially using tools. In the future, tools will elevate experience as their "alive" quality strengthens.

What ColaOS wants to achieve is exactly this — she wants to gradually adapt to users, letting each conversation build on the last, creating a sense of accumulation between you and her.

The year Her came out, everyone thought Samantha was science fiction. But in 2026, ColaOS has taken another step in that direction.

If you want to try it too, experience what it's like to be with an AI that "remembers you," that has Soul, you can go to colaos.ai and join her Waitlist.