August 3, 2026

The Office Agent Wars: Why Can't Big Tech Build a Good AI Product?

Big-company disease · The heaven-sent youth · The pre-install wars · Layoffs by another name
Contents

    “Your time, your head, your inspiration have already been drained dry by the items in your OKRs.” — Mark Tang, on why innovation doesn’t grow inside big tech

    On a Friday evening, on his way to get a massage, Raymond was lectured on how to work efficiently by the Focus Media screen in the massage building’s elevator — that being WorkBuddy, papered across every lift in China. China’s frontline tech war is all inside that elevator screen. DingTalk interviewees have to recruit five new users before reaching round two; Tencent’s model needed a “youth fallen from heaven” to untangle internal interests before it could ship; Feishu uses its own sibling’s model and gets not one cent of internal discount. Big tech has been all in on AI for three years — so why is everyone still holding the same products as before?

    In this episode, Raymond brings back EP40’s popular guest, Mark Tang, co-founder of Caidazi AI and host of the podcast Fācái Dāzi, and they name names across big tech at home and abroad: ByteDance’s Flow, Alibaba’s ATH, Google whose ad revenue rose rather than fell, Apple hitting record highs without spending on CapEx, the Doubao phone blocked by WeChat, the pre-install war that is certain to be re-run — and the question every working person cares about: has AI already started replacing people?

    What follows is the full conversation, edited and condensed.

    1. Why AI giants papered the elevators with ads

    Raymond: One Friday evening I was struggling to get to a massage, and the Focus Media screen in the massage building’s lift was running an ad: let me write your script, your slides, your diagrams, your report to your boss. I was extremely irritated — why, when I’m going for a massage, do you still need to remind me how to work? That’s the now very famous WorkBuddy. Focus Media is a witness to the front line of Chinese commercial warfare: this screen has previously told me to get free fried chicken for the World Cup, to use Afu for medical appointments, and how to get free bubble tea. China’s office agents have entered a new era, and we’re starting a new series called Every-Buddy — inspired by WorkBuddy and CodeBuddy. Who survives the next year and becomes the new super app? We have a returning guest who definitely knows the answer, Mark Tang.

    Mark Tang: Hello everyone, I’m Mark Tang from Fācái Dāzi. Last episode looked at where startups have got to; this one looks at big tech at home and abroad: what has progressed and what has stalled.

    Raymond: First question: have you been offered any Buddy sponsorship? I haven’t, and we need to reflect on why we’re this unsuccessful. An influencer friend of mine takes four different Buddy ads a month: he posts a dozen or twenty videos a month, so his ad load runs to twenty-odd percent, at a hundred-plus thousand yuan each — from Tencent alone he makes around 600,000 a month.

    Mark Tang: Zero ads — but we are available for purchase (laughs). Airport advertising used to be mostly cloud vendors like Baidu Cloud and Volcano Engine, and now it’s been handed down to applications. Focus Media amply proves that big tech has an absolute resource advantage. Getting you to see it in the lift every day and implanting the brand in your head costs a lot of money, and startups simply don’t have the resources. The previous wave was 2024, when Kimi and Doubao spent the most on performance marketing, and those were chatbot products.

    Raymond: There was a lot of criticism of Kimi at the time: all the money went to ByteDance, which amounts to giving ByteDance more ammunition to advertise Doubao itself.

    Mark Tang: People later worked it out: the more you spend, the higher the CPA, the customer acquisition cost, and the math gets worse and worse. Even today, chatbots haven’t found a suitable monetization model.

    2. What actually counts as an AI-native product

    Raymond: Last episode we covered startups like LibTV and MiniMax; today it’s big tech, but still within AI products. How do you define a good AI product?

    Mark Tang: Big tech has been all in on AI since at least 2023, and the wave swept nearly every large company —

    Raymond: Except Pinduoduo.

    Mark Tang: Pinduoduo is immune. But the products everyone actually uses can still be counted on one hand, and you’ve seen AI’s shadow inside all the old products: last year you could use AI to order food on Meituan; this year Alipay checks your orders and organizes your spending, and AI has appeared in WeChat — and none of it seems to catch on. I think building a product for a new era requires using the new era’s methods to redefine a problem only this technology can solve: something that couldn’t be solved before, and now AI can, or the efficiency goes to infinity. Beyond that, it can’t be called an AI-native product.

    Raymond: For instance in our investing work, many things previously couldn’t be done in bulk and at speed, and suddenly become possible — that holds up. But many products inside big tech are still “we could have done this before anyway.”

    Mark Tang: It just makes you a bit faster and a bit fancier. Many products forcibly convert a GUI, a graphical interface, into an LUI, a language interface, so finding a feature means going through a dialogue box, which is actually anti-human. There are a great many posts on Xiaohongshu saying “this product is doing AI for AI’s sake,” precisely because it wasn’t conceived from an AI-native standpoint. Some products have generated real volume: Doubao is basically a national-scale application, with monthly actives heading toward a billion like Gemini’s; WorkBuddy’s monthly actives should be over twenty or thirty million, and not just registrations pulled in by ad spend — there’s retention too. But fundamentally, nobody has seen the genuinely eye-opening super app of the AI era.

    Raymond: Public markets take the same view on Tencent: the share price has fallen all year and everyone is waiting for one especially impressive thing — don’t tell me WorkBuddy’s monthly actives went from 25 million to 30 million to 45 million. Just show me the thing; I’ll take one look and know it’s good, and the market cap goes up a hundred billion dollars immediately. Which is why everyone has high hopes for WeChat AI — Allen Zhang is the world’s greatest product manager, possibly without peer, and everyone is waiting for him to rescue Tencent’s share price. In the end WeChat AI launched after the Dragon Boat Festival with little community impact. So two questions: has anyone in the world pulled it off? And is big tech’s failure a matter of organization or of resources?

    Mark Tang: Let’s start with the challenges big tech faces. Caveat: my small-fry views certainly can’t match those of the strategy elites inside these companies.

    Raymond: Anyone is entitled to comment. My life principle is: I can’t refrigerate, but I can review a fridge.

    3. Why interviewing at DingTalk means first recruiting five new users

    Mark Tang: First is big-company disease. This year’s sensation was “Trapped Inside DingTalk”: one working stiff wrote an enormously long complaint post — an utterly inhumane overtime regime, an organization in total disarray, and a product that isn’t impressive either. And there’s the interview requiring you to recruit users from home — to interview at DingTalk you first have to bring in at least five new users at home, and in round two you’re asked to gather feedback from those five people and propose product requirements.

    Raymond: Is that real? A B2B-leaning product asking ordinary interviewees to go home and recruit users.

    Mark Tang: It’s real; there are activity metrics after all. That’s one form of big-company disease. Whether KPIs or OKRs, both are fundamentally task decomposition. OKRs’ advance was supposed to be bottom-up — “boss, what can I do to achieve your O” — but in practice goals are still set top-down, and a lot of the tasks handed down aren’t reasonable. There’s a term in organizations called local rationality: a big company has layer upon layer, tasks get infinitely subdivided, and the people below have extremely thin perception of the end user. The Q&A algorithm people are only responsible for raising answer accuracy and conversation turns to some number, and don’t care what the user actually perceives; same for product and growth operations, and legal and compliance too. Every department is doing the right thing as their own department defines it, and combined, the output diverges from the original goal. That’s the origin of big-company disease.

    Raymond: More people always means organizational efficiency losses — five thousand excellent people certainly don’t produce five thousand times one person’s output; the discount in between is very large.

    4. Why Tencent’s AI needed a youth fallen from heaven to rescue it

    Raymond: Let me add two observations. LatePost had a piece on Shunyu Yao going to Tencent to rescue Tencent’s AI (link in the show notes). What it describes isn’t how hard the model was to train, but how this person, through a series of fortunate circumstances, was able to continuously adjust his relationships with the leaders of various departments at Tencent so that everyone’s interests gradually converged; and the senior executive committee granted him enough trust and authority — like a magical youth fallen from heaven whom nobody on the committee envied or trampled — to give him the space and time to untangle internal interests, concentrate the force, and get the model from preview to V3 out the door. Training the model itself isn’t hard, and Tencent isn’t short of resources; there were simply too many hurdles in between. It’s the story of a large organization’s chronic ailments being removed one by one by an outside physician. Second example: Yubo mentioned on a podcast that while he was building Yuque, Alibaba was constantly cutting costs and improving efficiency, and Yuque, DingTalk Docs and Alibaba Docs coexisted as three departments. At first A was going to swallow B and B was going to swallow C, and then none of them swallowed anyone. Then it was said Alibaba Docs should merge Yuque and DingTalk Docs, and Wuzhao disagreed — so in the end there was no merger. Yubo concluded building this inside a big company was impossible, and eventually left Alibaba to found his own company.

    Mark Tang: It’s all the same: you have to persuade too many people inside a big company and move too many people’s slice of cake, so you may as well go out and do it yourself. A company doesn’t hand you resources just because you have an idea: you go through an innovation project review, and the deck has to argue not just the idea itself but what data it will bring, how many users you expect, how big the market size is, where commercial growth comes from, how you’ll acquire customers — far too many questions to answer at a very early stage, so the drive to innovate is very weak. A common big-company approach to innovation is to give a team ample resource allocation and let them get on with it without reporting milestones — relying on a small team’s agility, effectively internal incubation. Google’s NotebookLM was hacked together by a few people in their spare time, initially from an “AI turns things into podcasts” idea, not a top-down directive that Gemini should build a note-taking system.

    Raymond: But in practice with internal ventures, HR will ask: who’s allowed to found something? For how long? Does it count toward performance? How are the options valued? Ultimately the big boss still has to rule on it. Luo Fuli — Xiaomi’s large-model lead and formerly DeepSeek’s model lead — said on a podcast that Lei Jun gives them enormous room and doesn’t come around asking how the numbers look; report it yourself when you have a conclusion, and the rest is up to the team’s own drive. That kind of thing encourages innovation.

    Mark Tang: But at companies like ByteDance and Alibaba, once a small team produces results it enters horse racing. People’s heads — product managers’ especially — are roughly alike: once a direction is clearly defined, the solutions anyone can come up with are all similar, since not everyone is Steve Jobs. Everyone sees the same problem and thinks of roughly the same approach, and in the end the leaders confer and merge mine away — horse racing is, in a sense, precisely not encouraging innovation. I believe Yubo had thoughts like this when he left to build YouMind. And there’s the innovator’s dilemma: a mature company’s legacy business is extremely profitable — Tmall is still the revenue mainstay, Douyin is a money printer —

    Raymond: Douyin runs harder than a money printer. To actually print that much money the printer would occupy an enormous factory floor and the annual electricity bill would be substantial.

    Mark Tang: So having me build something new is less motivating, because I’m being asked to lift a rock and drop it on my own foot — the new business is very likely there to replace my old one. Even knowing the rock might smash something beautiful out of the ground, you can’t bring yourself to drop it.

    Raymond: To summarize: big tech’s existing business forms are extremely stable, and forcibly improving them with AI is sometimes backwards, like a GUI that shouldn’t be forced into a dialogue box; and even when a real idea is found, they hesitate and don’t dare execute — which is predicated on their being capable of finding it at all, because it has to go into the OKRs.

    Mark Tang: Even for a side project, a big-company worker’s time is entirely filled, and it’s actually false busyness. Your time, your head, your inspiration have already been drained dry by the items in your OKRs.

    5. Why ByteDance’s AI runs faster than anyone’s

    Raymond: This set of problems — the innovator’s dilemma, organizational disease — should theoretically afflict Alibaba, Tencent, Baidu and ByteDance alike. So why does ByteDance seem to do better, with Tencent and Alibaba in the middle and Baidu worse? Xiaomi has actually moved forward a bit because its tolerance on models is high enough.

    Mark Tang: The management models differ. ByteDance established the Flow department in 2024, led by Musical.ly founder Alex Zhu — a product manager by background, and the company wanted him to design the department from a product standpoint. Under Flow sits Seed, where all of ByteDance’s large models live; Seedance and Seedream both came out of it. Crucially, it’s completely separated from Douyin, Fanqie and Toutiao — it isn’t an appendage of some old business. Many failed approaches consist of adding AI features into an old product: add it into Alipay, add it into WeChat, and it’s the same old team doing it, which brings you back to lifting a rock onto your own foot. ByteDance established an entirely new department, running relatively independently since 2023 and 2024, with independently priced independent options; Maoxiang also came out of Flow. Independence guarantees tolerance, and it doesn’t have to carry too many short-term metrics. Flow also doesn’t manage Feishu; that’s completely separate.

    Raymond: So ByteDance can be viewed as two lines, Flow and Feishu. Feishu has assembled every big-company malady we just described — so why did it also break out?

    Mark Tang: Feishu’s organizational DNA is unusual: it started as a ten-person internal project building ByteDance’s internal IM, with no KPIs and no commercial growth pressure. It’s now several thousand people and has been through a very large round of layoffs, but the DNA is still “solve our own problem first,” while simultaneously productizing all of ByteDance’s advanced organizational management practice — communication mechanisms, document systems, review mechanisms, legal and compliance — and packing it into one product. Only by having lived through those things inside a big company can you design a communication system suited to an excellent organization. And Feishu and Flow aren’t in a tightly bound cooperation; it’s an internal settlement relationship: Feishu, or Lark, has to pay for using models on Volcano Engine, and internal accounting gives no discount — it’s more expensive than the finished product, so you have to find ways to polish your own thing more finely. Conversely, feedback data can also be settled the other way — the traces of the working stiffs who use Feishu can be sold as data packages to the model side, to Volcano. Even blood brothers settle accounts clearly, which beats two teams coupled together muddling through on fuzzy books.

    Raymond: So why is ByteDance better: the department running AI is relatively independent and very loosely coupled with the traditional business, and the person running it is the big boss’s most trusted, relatively young. What about Alibaba?

    Mark Tang: Alibaba established ATH this year — Alibaba Token Hub: production of tokens, supply, and pushing them down into applications, all inside this one business group, with the boss being Eddie Wu himself, the CEO directly in charge.

    Raymond: AI is certainly a number-one-executive project — pick anyone else and it’s hard to command respect and hard to marshal resources. ByteDance and Alibaba, two organizations that can mobilize tens of thousands of people, adopted two different organizational forms in the face of AI. What else does ByteDance do well?

    Mark Tang: Long-standing organizational culture. Many people complain about the “Alibaba flavor” — the jargon of reporting upward. ByteDance has less of it, though certainly some; grab the business pain point, find the growth handle, get a big result, align on granularity — everyone says these. That relates to ByteDance’s hiring preferences: hire cleverer, more self-driven people who can find their own way out if locked in a room. People used to say that ByteDance’s Feishu building and Alibaba’s DingTalk building are next door to each other and they compete over whose lights go off later at night — Feishu may have left the lights on with nobody inside, burning up DingTalk employees’ sleep. But DingTalk simply doesn’t leave; that’s been the corporate culture since Wuzhao’s time. I think it’s appalling.

    6. Nobody uses Google search anymore, so why is ad revenue still rising

    Raymond: Having gossiped about Chinese big tech, let’s do overseas — defined as the Magnificent Seven type of company that can hire tens of thousands.

    Mark Tang: Google first, one of the relatively more successful big companies on AI applications. Everyone kept fearing search would be replaced by AI and hit ad revenue, and the last several quarters have falsified that: ad revenue not only didn’t fall, it rose considerably.

    Raymond: Let me interject something I haven’t figured out: I genuinely never use Google to find data or answers anymore — I still use search, but only to find a particular image I’ve seen, or a particular web page, and the depth and volume of queries has absolutely dropped substantially. Why is advertising still rising? Who filled in my gap?

    Mark Tang: The business just doesn’t drop, which is remarkable — who exactly is using it? Unknown. But search’s share hasn’t fallen as much as everyone predicted, Gemini’s actives keep growing, and on the web it has taken some volume from ChatGPT. When Gemini 3 came out late last year, Google was called “the god of AI”: front-end ability is especially strong, and there were vast numbers of people on Xiaohongshu using it to draw product prototypes. Internally there’s a whole suite too: Antigravity for writing code, AI Lab for interaction prototypes, Stitch for UI/UX.

    Raymond: But it’s quieter this year and the base model isn’t outstanding. Is the application’s diminished volume related to the base model’s diminished volume?

    Mark Tang: It’s related. A lot of people switch between Codex and Claude Code purely on whose model is newest and strongest — loyalty to a model is zero; you use whoever is stronger. I do think memory is a moat, but agent products’ memory is all local right now: I move from Claude Code to Codex and it automatically migrates local memory, MCP and skills over in one click, painlessly.

    Raymond: Domestic big tech’s work-type products are different again: they all connect to open-source models and you pick. Nobody uses WorkBuddy because “I just love Hunyuan”; once inside, they still pick DeepSeek.

    Mark Tang: That’s a short-term artifact. Suppose WorkBuddy establishes itself, reaches 50 million daily actives, essentially saturates China’s office population, and Hunyuan 3 has entered the domestic first tier — why would it still let other open-source models plug in? It can enclose the whole thing and form a more closed ecosystem. It’s open in the short term because my model isn’t the best but I want my harness to be better; once the model gap is minimal and users won’t abandon my harness over it, I have a lot of room to tighten up.

    Raymond: But many products’ marketing follows the model cadence — “we’re the first video editing product to integrate Seedance 2.5.”

    Mark Tang: That’s LibTV’s approach; WorkBuddy may ultimately follow Hunyuan’s own cadence.

    Raymond: I’d argue it has no reason to shut out open-source models. A query I type into WorkBuddy, even if it’s routed to GLM-5.2, doesn’t reach Zhipu — Zhipu doesn’t get the input and output and can’t use WorkBuddy’s data for post-training. Whereas whatever model you use, Hunyuan receives your input and output and also learns how GLM-5.2 calls tools. The whole data loop sits with the operating vendor, so it has no reason to shut them out — unless one day it wants to compete on the model’s public profile.

    Mark Tang: Put that way you’re right, in which case it would have to make the model closed-source and take the top spot on the leaderboards — if you aren’t SOTA, then you’re just hawking your own melons.

    7. Why WeChat AI always has a chance

    Raymond: Besides Google, which other overseas giants are doing well?

    Mark Tang: Very few. Meta? I previously commented on their renting out compute: it doesn’t mean the industry has surplus compute, but it absolutely means Meta itself has surplus compute — the industry is still climbing and selling it carries a premium, but their own applications haven’t taken off. The Manus acquisition is the same: Manus keeps getting better and Meta doesn’t seem to have learned anything from it; the diffusion of general-purpose agents is obvious, everyone has learned it, and in America the competitors are too strong — two of them are enough to handle the whole thing. Meta and WeChat are actually in the same race: Facebook and WhatsApp both have enormous potential. I remain firmly convinced WeChat AI can be an extremely good product; its context is too distinctive — whether private one-to-one or group chat, there’s so much it could do, as long as it gets the measure right. Same for WhatsApp, and Telegram hasn’t done anything either. What they’re doing now is very like Feishu: add AI as a member of a group chat, like Slack — but without drawing on your relationship graph or chat history. It may be privacy considerations, or it may be that they don’t know how to strike the very narrow line of “let you feel the benefit without feeling violated.” But given how long everyone spends chatting and handling things in IM every day, it’s certain that an extremely good product can grow there.

    Raymond: Give them more time. At minimum this isn’t a startup’s opportunity.

    Mark Tang: There’s an American company called Poke — not Bill Zhu’s Pokee, a different one, apparently also with a Chinese founder, building an IM assistant — and its design thinking is good: give it two authorizations, email and calendar, and it can cold start. It’s a personal assistant with a touch of companionship: from those two authorizations it knows what you’re doing and what needs you’ve expressed, proactively helps solve problems, and then messages you over iMessage — riding on a traditional product form while using AI to do the job better. This company is also about to be acquired, which shows people are watching this space. Wherever user attention is abundant, flowers naturally bloom.

    8. Why nobody doing AI buys a Windows machine

    Raymond: Let me run the menu: Gemini is good, Meta a notch below. What about Microsoft and Apple?

    Mark Tang: Microsoft is probably just in a hard spot. I can’t picture its opportunity on the application side; Azure’s cloud will be a very important part, and fundamentally it pivots to infrastructure.

    Raymond: There was news the other day that Microsoft might acquire Mistral. But when tool calling and vibe coding first appeared, everyone had enormous hopes for Microsoft: you are, after all, the productivity tool for everybody on Earth.

    Mark Tang: First, Windows is unfriendly in the AI era: if you’re doing AI, basically nobody buys a Windows machine. Mac has unified memory — 24GB or 36GB shared across the whole system, unlike Windows partitioning workspaces; the graphics are integrated, the SSD is integrated, and you don’t have to shuttle data around. All the hardware is soldered together.

    Raymond: Same logic as how we now look at AI hardware: originally GPU and DRAM were designed separately, now the GPU and HBM are soldered together, Vera Rubin puts CPU and GPU together, and future advanced packaging will pack SSD and DRAM in as well.

    Mark Tang: Reducing the loss from information transfer gives the best performance and the best energy profile: the new MacBook Pro can run twenty hours on light office work, which Windows basically can’t. Apple’s push on M-series chips is visible to everyone, and the sales share is obvious; new AI applications all support Mac chips first. Mac isn’t yet the overwhelming majority among knowledge workers, but it’s switching over, and the potential is very large — even last year’s cut-down-chip MacBook Neo sold especially well. Second is the Office suite: why will we still need slide decks in future? The form of Excel and Word persists — people simply need a form where cells add up to a number — but it doesn’t have to be Microsoft Excel: Google Docs and Feishu Docs are both especially well built and come with many additional features; Copilot hasn’t kept up, so I may as well use a product that wraps in more mature AI capability.

    Raymond: I’ve used Mac all these years and have never installed Excel. So Microsoft’s most valuable part is that it obtained equity in OpenAI; the infrastructure is good and it does FDE well — helping more enterprises connect to AI and do organizational transformation, which is also very valuable.

    9. Apple doesn’t spend on AI, so why is it at a record high

    Raymond: And Apple?

    Mark Tang: Apple is very interesting: from early on to today it has barely participated in the AI war, with no enormous CapEx, and it should now be the most valuable company in the world, with the share price at a new high.

    Raymond: The market used to reward whoever spent CapEx, and has now started punishing it, entering a phase of rewarding not spending CapEx — Apple isn’t chasing data centers and isn’t building a cloud.

    Mark Tang: Its hardware ecosystem advantage is extremely strong, share is expanding, chip processes keep improving, and it has the ability to pass costs on to consumers — becoming the McDonald’s of consumer electronics: every price increase gets passed to users and they still want to buy, which proves the scarcity value is there. Many people bought their computer after the price rise. Apple hasn’t produced especially outstanding hardware these past few years — Vision Pro was fairly unsuccessful, high price and low volume, and AR is gradually weakening — but it still has a chance, starting with Apple Intelligence. Same logic as WeChat: people use it every day, so you have a chance. Since 2024 it has shipped a lot of protocol-level things: Caidazi as an app plugs into Apple, and there’s a whole set of components letting Apple Intelligence, as the OS, invoke your application — the same logic as WeChat giving mini-program developers an SDK so that WeChat, as the OS, can call mini-program interfaces.

    Raymond: Does Apple Intelligence mean Siri? Where’s the entry point?

    Mark Tang: Siri is part of it, and the entry point is your phone: long-press to invoke, hand it a task, and when it finds it can’t solve it itself, it searches whether an app you’ve installed happens to be able to, and invokes it for you. For instance you say “I want to get rich” —

    Raymond: It traverses all my apps, finds Bloomberg and the Financial Times and something called Fācái Dāzi, and pulls it up.

    Mark Tang: Right, same as WeChat AI: you say order some food, it pulls up the Meituan mini-program to order for you, and finally gives you an entry point you tap into to see the GUI with the order completed. Two levels: phone to app, app to mini-program.

    Raymond: The mini-program system basically only exists on WeChat and Alipay, doesn’t it? What about ByteDance?

    Mark Tang: Douyin has them too, with lower penetration. But ByteDance has done local services for so many years — Dianping is in real jeopardy — that Douyin covers every scenario you can think of; it just hasn’t reached WeChat and Alipay’s level.

    10. Why the Doubao phone got blocked by WeChat

    Raymond: This is a good moment to discuss the Doubao phone — is it a concept, or what state is it in?

    Mark Tang: The Doubao phone is an OS-level matter. You tell it “find me what Raymond talked about in the first episode of Mossfire,” and it opens your Xiaoyuzhou app, searches Mossfire, finds the episode and pages down through it — that’s computer use. It was previously blocked by apps like WeChat for exactly this kind of operation and got paused, but it will very likely return.

    Raymond: People say the next entry point is the opportunity at the device.

    Mark Tang: Everyone wants to seize it — Apple is being rewarded precisely for doing the device well. Domestically OPPO, Huawei and vivo are all building AI-native phones: OPPO’s recent foldable sells for nearly ten thousand yuan, and the marketing slogan is that it comes pre-installed with a lot of AI-native things, pre-installed at the OS layer with a local model, the same thinking as Huawei’s Celia. There are two approaches for the OS operating apps: computer use — screenshot to the model to understand, tap what should be tapped, swipe what should be swiped, a machine forcibly dressing up as a human doing things; or the app exposes interfaces, MCP interfaces, and the OS calls the API directly. The latter tests how the OS reaches agreements with that many apps — the app layer has to open its capabilities up.

    Raymond: For instance I tell my phone I want roast duck, it calls Meituan delivery, fills in the address, and thirty minutes later the duck arrives, with no need to open Meituan’s GUI at all.

    Mark Tang: That basically all works now; Alipay and WeChat AI can both handle ordering food, and all that’s left is you pressing to confirm payment.

    Raymond: Let’s use this to clarify the concepts: what’s the difference between an OS agent and a super app agent? WeChat and Alipay are super apps, but they aren’t super app agents, and they aren’t OS agents.

    Mark Tang: It depends who the user ultimately interacts with. An OS agent is you talking to the hardware, and the hardware dispatching the task to apps — Siri is closest to that state, and domestic Android phones are like car systems, “Li Xiang Tongxue,” “Xiaoyi Xiaoyi.” A super app agent is you talking to an app, and the app dispatching the task to mini-programs inside it. This fight will be fought: when WeChat built mini-programs it wanted to become a quasi-OS, and it’s even capable of building a WeChat phone. How it’s fought depends on two factors. First, experience: talking to WeChat AI requires opening WeChat first, inherently one extra step; talking to the hardware can be invoked with the screen off — on this dimension hardware is stronger.

    Raymond: But if there’s a very clear killer app — WeChat — that won’t grant permission and won’t integrate, then this hardware agent becomes you can do everything except WeChat, which is also painful.

    Mark Tang: Then it comes down to whether hardware vendors can get this computer-use visual operation working — it only needs screen permission, the same as screen sharing on a Mac.

    Raymond: So I buy a ByteDance phone and install ten apps: five grant permission and are governed by the OS’s agent; the other five — Meituan delivery, WeChat, Alipay, Didi, Ctrip — don’t open up. ByteDance says never mind, I’ll turn on computer use and achieve the same end, just slower, so I’ll improve chip performance. You can’t ban computer use. So it sounds like the OS agent wins more easily?

    11. Will Meituan get compressed into a kitchen

    Mark Tang: There’s one more thing: who ends up with the user’s conversational data. There are three players on the field: the OS, the app, and the mini-program builders. My earlier view was: in future many vertical applications may exist as subagents and never reach the entry-point position — nobody opens the Caidazi app specifically to ask one investing question, but a WeChat conversation can call Caidazi’s Q&A entry and hand the answer to the user. So what data does each of the three get? You tell the OS: order me some malatang, then get me a cab to 12 Nautical Miles. The task splits and goes to Meituan and Didi — they work better as subagents in their own domains, but Didi doesn’t know you have a habit of eating malatang —

    Raymond: And Meituan doesn’t know I like coming here to record podcasts.

    Mark Tang: So your personal context and memory get lost in the middle. Long term, whoever the user ultimately talks to is the one holding the most complete picture — whoever can remember what kind of person you actually are has the whole of it; everything below is a dispatched role receiving only a partial task. The most Caidazi can obtain is your investing questions, with no idea that Raymond likes malatang or where he shops every day.

    Raymond: Context is extremely valuable — the reason I as a person can be a mine is that it holds my context. So in two years, will everyone start heavily subsidizing phones and devices? Use my phone and I’ve got you under control; and there’s a very high probability models become small and refined and run directly on device.

    Mark Tang: That holds. You don’t even need on-device small models: even if on-device capability is insufficient you can connect an open-source model in the cloud, and the context still sits with the OS and is still captured.

    Raymond: Last episode I discussed this with Wang Tiezhen: everything I do in Claude Code gets taken to train new models and I’m already at peace with it — just as, having mentioned malatang this many times, I guarantee the first page of Meituan delivery will have malatang. Mr Wang said one shouldn’t be at peace, and talked about installing models locally and the open-source community, and also raised the point that protecting privacy and not selling privacy can itself be a business — in which case Apple might be the biggest beneficiary.

    Mark Tang: Looked at the other way, big tech ultimately all wants to establish a direct connection with users rather than be treated as a tool by a unified hardware dispatcher. Right now if you order Meituan delivery through WeChat AI, Meituan still gets the user ID and knows it’s sending malatang to this gentleman Raymond; in future WeChat won’t pass the user ID and payment won’t happen at Meituan, and what Meituan receives is one instruction: deliver this item to this address. That’s it. Recommendation algorithms and ranking become meaningless — it has no contact with the user and is purely the dispatched party.

    Raymond: Compressing Meituan into the role of a kitchen.

    Mark Tang: So big tech all needs to move up the stack, and certainly can’t accept that. Also: pre-installed apps on Android will intensify, and the narrative is about to make a comeback —

    Raymond: Live long enough and you get to experience the same thing twice. The previous pre-installed apps were for people to use.

    Mark Tang: It used to be that opening a Xiaomi phone gave you certain things, and Apple’s default browser for American users is Google — all of which commands a price. Now it’s: as an app, I want the OS’s agent to invoke me more, so I buy pre-installation, and even sign an exclusivity agreement — intents of this type can only be routed to me. That could well happen. Because building and shipping an app now means supply has become infinite supply — anyone can hack together an extremely good application with Codex or Claude Code. In the end it becomes a fight over channels and a fight over traffic, the same logic as phone pre-installation.

    Raymond: So the next few years may re-run a pre-install war: ByteDance says please install my office software and don’t install WorkBuddy; a shop assistant at some Android phone store takes a kickback and, the instant the user powers on, says “let me install a few useful apps for you,” and installs eight office suites — which happened in China between 2010 and 2013 or 2014.

    Mark Tang: Exactly, and here we go again.

    Raymond: So what stock should we buy? Sunny Optical? Any good public-market names, listeners are welcome to discuss (laughs).

    12. How Doubao actually evaluates whether a product is good internally

    Raymond: It suddenly occurs to me: is big tech’s internal product evaluation all just daily and monthly actives?

    Mark Tang: Certainly not. Doubao, for instance, staffs at least twenty or thirty product managers doing nothing but evaluation — they do nothing else, defining evaluation standards and scanning queries and looking at results all day. They ultimately have to resolve into outcome metrics like daily actives, monthly actives and monetization, but upstream you decompose into process metrics: average conversation turns; whether there’s negative emotion within a multi-turn session — “how can you get even this wrong,” “wrong again, didn’t I just tell you above?” — those are red-light signals; and retention — you asked in the morning, do you ask again in the afternoon, and given an expectation of how many questions you have per day, how many route through to my application.

    Raymond: That sounds no fundamentally different from traditional internet product evaluation?

    Mark Tang: It isn’t, just like e-commerce from impression to click to order to delivery confirmation; only the weights have changed: it used to be that I absolutely had to get you to spend enormous amounts of time on me, and now you can spend hardly any time as long as you pay. Monetization becomes a very important metric — you can serve just ten thousand people at ten thousand yuan each and have a hundred million in revenue, without needing to serve a hundred million people at ten cents each.

    13. Has AI already started replacing people

    Raymond: Last question: how do you see the endgame of the office agent war? How do you frame the TAM of Chinese white-collar workers who will use AI? At what magnitude, at what revenue metric, does it end? How does the next twelve months develop?

    Mark Tang: Being asked to make predictions right now is genuinely hard.

    Raymond: It’s fine, listeners just listen today and won’t come back in a year (laughs).

    Mark Tang: Within twelve months it’ll certainly still be competition centered on office scenarios, because that’s where the value is greatest: programmers have the highest salaries; below them are white-collar written work in legal and finance; below that people writing copy, with declining value. What’s being fought over right now is that portion of the market, with a large enough population and enough GDP generated. Fundamentally, AI is here to replace people; it needs to produce more GDP — it’s just that the replacement progress bar hasn’t yet reached the point of generating additional value.

    Raymond: So we’re still at enablement, not replacement? I think within 12 months we certainly reach replacement — Fable is already there, and I’m already on the edge of replacement.

    Mark Tang: Right. The China-US model gap used to be 18 months and is now under six to nine, possibly within half a year: GLM-5.3, about to ship, and Qwen’s new model are basically at Opus’s level.

    Raymond: Our earlier judgement was that Opus 4.8 is a sufficiently usable model — the “usable” line has been passed, so it’s completely beyond enablement and on the road to replacing people.

    Mark Tang: Which is why layoffs are already genuinely happening. China hasn’t been especially fierce — people do have to have jobs, after all; America is extremely dramatic, with Microsoft, Meta and Amazon all cutting heavily.

    Raymond: And all at 10%. But is that fierce? The absolute headcount is large, but given that every American can use Fable, why cut only this much?

    Mark Tang: It will accelerate; cutting is addictive, and the more you cut the more you cut. Right now humans still have to play reviewer and code checker, and as those jobs also get gradually replaced, more people become unnecessary. Fresh graduates especially don’t know what they can do — one old hand with five agents is far more formidable than your five old hands with five fresh graduates.

    Raymond: Are Chinese layoffs more covert? Tencent isn’t going to announce cutting a hundred thousand people today; Silicon Valley giants, by contrast, put out very aggressive layoff plans to please shareholders and the stock rises on the spot.

    Mark Tang: As I understand it, it’s layoffs by another name: previously mobility between big companies was extremely high, people changed employers after six months or a year, and a team losing two people backfilled two; now maybe five leave and one gets backfilled — which is layoffs by another name.

    Raymond: Headcount is something to watch in the financials going forward.

    Mark Tang: Of course, a company like Alibaba has many sales and B2B roles, so fluctuations don’t entirely represent white collar being replaced. In the first half of this year the ones cutting first were tech companies, because knowledge workers’ assets can be digitized; they’re all in the computer. Once the data has been sampled and distilled, the person is no longer needed. A very sad thing. This will gradually extend into traditional industries: with this many companies doing FDE now, non-tech industries will also bring in AI processes and tools, and what employees do gets scanned in and distilled away too — work that used to take three people, one person now does with three AIs. And it’s a top-executive project: you can’t get anywhere with the product director or the R&D director, but with the CEO you get there immediately — his arithmetic is crystal clear.

    Raymond: Only the person paying the salaries feels it.

    Mark Tang: Silicon Valley also went through a phase of token maxing: the company says AI, go explore quickly, everyone pick up the pace. Now the contest is over — not that token maxing has no effect, but you have to have output. Whoever genuinely got more efficient, I can see it; whoever runs a pile of agents all day claiming they burned tens of billions of tokens with no results to show — are you just playing around? A company won’t let you burn tokens infinitely.

    Raymond: So token maxing is a timed contest: fixed time, unlimited resources, see who can amplify output, and keep the people who can.

    Mark Tang: If you can’t amplify output you’re wasting money — I give you the most advanced tools, and if you produce good results that’s you being good and I reward you; if you can’t produce, why keep you.

    Raymond: So how many people do you think Tencent should keep in the end?

    Mark Tang: No idea; you’d have to ask Mr Ma.

    Raymond: At the end of 2022 Musk acquired Twitter, walked into headquarters carrying a sink, and cut 80% of the staff that day, from a few thousand to a few hundred. Everyone then discovered Twitter turns out will be fine — and that was before GPT-3.5 had even shipped. If he acquired it again today he might cut all 100 remaining: all of you leave, I’ll do it myself.

    Mark Tang: It wouldn’t come to that. You always need some people — which is why the OPC, the one person company, doesn’t quite hold up; at minimum it has to be a three-person company. But it’s genuinely become a bit more feasible.


    Recommended by our guest: Mark Tang’s podcast, Fācái Dāzi.

    Full Video
    Watch the full episode on YouTube →

    If you're working on this too — or you think we've got it wrong — write to us at [email protected]; if you'd rather not write, just leave your email below.

    Topics AI ApplicationsWork & Organizations
    Disclaimer. This content is for general informational purposes only. It does not constitute investment, legal, tax, or accounting advice, nor an offer to sell or a solicitation of an offer to buy any security or interest in any fund managed by Mossfire Capital. The Firm and its affiliates may hold positions in the instruments discussed; actual positions may differ from — and even be contrary to — the views expressed, and are not disclosed. See our full Disclosures.