
A Wisdomia White Paper: Part I — The Format and the Market: Anatomy of a Vertical Revolution
Seventy-three thousand years ago, in a cave at Blombos on the southern Cape of Africa, a human hand ground ochre into a paste and scored a lattice of crossed lines onto a piece of stone. There was no market for it. No audience metric. No retention curve. Only the irreducible human compulsion to compress the world into a symbol and hand that symbol to someone else.
Everything that follows in this paper, every LoRA weight, every hook-rate benchmark, every coin bundle priced at £4.99, is a downstream tributary of that single gesture.
We are now watching the strangest chapter of that long story. Storytelling, the oldest human technology, has collided with generative artificial intelligence, the newest, and the collision point is not a cinema, a broadcast tower or a streaming service. It is a rectangle roughly the size of a human palm, held vertically, watched with one thumb poised to scroll.
The micro-drama, duanju in Mandarin, vertical drama in the trade press, is a serialised scripted narrative of sixty to ninety seconds per episode, shot 9:16, built for the phone and only the phone. It looks, to the untrained eye, like the most disposable artefact our civilisation has yet produced. It is not. It is a precision instrument, tuned to the exact frequency of human attention, and it is currently earning more money in China than the country's entire domestic box office.
What is actually happening is this: Scheherazade has been industrialised.
In One Thousand and One Nights, a woman survives by refusing to finish a story. Each dawn she breaks off mid-sentence, and the king, who intended to kill her, postpones the execution to hear the ending. The cliffhanger is not a narrative device in that book. It is a survival mechanism. It is the original paywall.
Dickens understood it, publishing The Pickwick Papers in monthly shilling parts and leaving readers hanging on the docks of New York waiting for the ship carrying the next instalment of The Old Curiosity Shop. Aristotle had already named the machinery in the Poetics: peripeteia, the sudden reversal; anagnorisis, the moment of recognition. The micro-drama does nothing Aristotle did not describe. It simply does it every fifty-eight seconds, in nine languages, generated by a model, and charges you a coin to continue.
That is the subject of this arrticle. Not whether it is art, a question that has never once been settled in advance for any new medium but how it works, what it costs, who is building it, what the law is doing about it, and what it tells us about the shape of human attention in the age of machines.


A micro-drama is a serialised scripted narrative designed exclusively for vertical, mobile-first consumption. Its defining constraints are structural rather than thematic:
Attribute | Specification |
| Aspect ratio | 9:16, vertical, single-thumb operation |
| Episode length | 60–120 seconds; increasingly compressing toward 60 |
| Series length | 60–100 episodes (Western apps), 22–42 episodes (AI-native pilots) |
| Total runtime | 80–120 minutes — the length of a feature, atomised |
| Narrative unit | One reversal per episode, one hook per ending |
| Cast density | Rarely more than two faces in frame; the phone cannot hold a crowd |
| Monetisation | Free episodes 1–15; coin unlocks, rewarded ads or subscription thereafter |
The genre repertoire is deliberately narrow and deliberately universal: billionaire romance, secret heir, revenge on the family who wronged you, werewolf and vampire fantasy, time-travel redress, the wronged wife who returns transformed. These are not new. They are the load-bearing walls of folk narrative, Cinderella, the Count of Monte Cristo, the Ugly Duckling, the returning king, rebuilt for a device that fits in a pocket.

An AI micro-drama is one in which generative models materially perform or assist the work of production: script generation, storyboarding, scene composition, synthetic performers, facial expression transfer, voice cloning, dubbing and localisation. Three tiers have emerged in practice:
The distinction matters enormously for law, for labour and for cost and we return to each in Parts II and III.
The format did not emerge from film. It emerged from the collision of three Chinese ecosystems: web-novel publishing, mobile gaming monetisation and short-video distribution.
Web-novel platforms, China Literature, Tomato Novel, had already trained hundreds of millions of readers on serialised, chapter-locked, pay-to-continue narrative. Mobile gaming had perfected the coin economy. Douyin, WeChat Video Accounts and Kuaishou provided distribution with payment rails already embedded. Micro-drama is what you get when those three systems are wired together and pointed at melodrama.
Chen Bo, chief executive of Neorigin, has traced the format's evolution directly from gaming methodologies and web novels rather than from television. That lineage explains almost everything about how the business behaves: it commissions like a game studio, tests like a performance marketer and distributes like a social platform.

The scale achieved is difficult to overstate. Revenues moved from $0.5 billion in 2021 to $9.4 billion in 2025, an increase of roughly nineteen times in four years, and in doing so overtook the entire Chinese theatrical box office. By 2030, advertising is forecast to contribute 56 per cent of Chinese micro-drama revenues, subscriptions 39 per cent and commerce 5 per cent.
Vivek Couto, executive director of Media Partners Asia, delivered the sector's most quoted verdict at the point the numbers became undeniable: "It's no longer a fad." He characterises micro-drama as a new entertainment and monetisation layer sitting between social media and streaming.
That framing, a layer, not a genre, is the correct one, and it is why incumbents were slow to see it. Micro-drama did not attack television. It occupied the interstitial minutes that television never had a product for: the queue, the commute, the ten minutes before sleep.

| Metric | Value |
| Global revenues, 2025 | ~$11bn |
| Global revenues, 2030 (forecast) | ~$26bn |
| China, 2021 → 2024 → 2025 | $0.5bn → $7bn → $9.4bn |
| China, 2030 (forecast) | $16.2bn (11.5% CAGR) |
| Ex-China, 2024 → 2030 | $1.4bn → $9.5bn (28.4% CAGR) |
| United States, 2024 → 2030 | $819m → $3.8bn |
| Japan, 2030 (forecast) | >$1.2bn |
| Global viewers | >830m, ~60% paying |
| Cumulative category downloads | >2.3bn |
| Ex-China quarterly revenue, Q3 2025 | $800m (2× Q3 2024) |
Three structural facts deserve emphasis.
First, the growth is now outside China. A 28.4 per cent CAGR ex-China against 11.5 per cent inside it means the centre of gravity is migrating. China remains the largest single market and the source of the operating playbook, but the incremental dollar is increasingly American, Japanese, Korean, Indian and Latin American.
Second, the cost structure is inverted relative to traditional media. Couto's formulation is precise: production is cheap, distribution is expensive, and success depends on speed, scale, and repeatable IP. In cinema, the negative cost dominates. Here, customer acquisition cost dominates. This single inversion explains why micro-drama companies raise user-acquisition debt rather than slate finance, a point we quantify in Part III.
Third, profitability is demonstrated. DramaBox's $323 million revenue and $10 million net profit in 2024 is the proof point the category needed. Hernan Lopez of Owl & Co put the discipline plainly when assessing the new financing structures: "You have to prove the unit economics are there."

Micro-drama does not travel uniformly. It travels along the grooves cut by payment infrastructure, gender demographics and existing serialised-reading cultures.
United States. The most lucrative ex-China market. Adoption is driven by affluent urban women aged 30–60, with romance, CEO storylines and revenge narratives dominating. DramaBox and ReelShort lead the app charts; DramaBox's participation in the Disney Accelerator signals major studio attention.
Japan. Forecast to become the largest APAC market outside China at over $1.2 billion by 2030, supported by LINE Pay integration and growing local production capacity. Japan's manga-and-serial reading culture is the deepest pre-existing substrate the format has found anywhere.
South Korea. Where AI-native production is furthest advanced commercially. Vigloo, operated by SpoonLabs and backed by an $86 million stake from games company Krafton, is applying K-drama craft to the vertical format. Krafton's thesis is instructive: the serial engagement loop of vertical drama resembles the engagement loop of a live-service mobile game closely enough that the same audience-building and monetisation logic transfers.
India. Officially still "exploratory" in MPA's assessment and simultaneously the most interesting laboratory in the world for AI-native production, which we examine in detail in Part III.
South-East Asia and Latin America. Identified as the strongest emerging growth territories. MPA analyst Myat Pan Phyu singles out Thailand for a 360-degree model distributing through both OTT apps and mobile networks with dual ad-and-subscription monetisation.
Africa. Barely measured, and therefore the most under-priced opportunity in the category. Adrian Cheng, chair of Crisp, described micro-drama as a phenomenon running from Seoul to São Paulo, noting that in Lagos comedians already thrive in two-to-three-minute vertical scripted video, and arguing that every frame is intentional, every moment carries intention.

The micro-drama could not have existed in 2015. Four curves had to cross.
1. The device became the theatre. Vertical video was a joke until it was 70 per cent of consumption. The phone did not adapt to cinema's grammar; cinema's grammar was discarded.
2. Payment rails collapsed into content. WeChat mini-programmes, in-app coin purchases, LINE Pay, rewarded video — the distance between "I want to know what happens" and "I have paid to know what happens" fell to roughly one second and one thumb-press. Desire and transaction became adjacent. That adjacency is the business model.
3. Attention restructured itself. I have argued in The 5th Industrial Revolution that the defining scarcity of our era is not compute, capital or even energy, but coherent human attention. Micro-drama is the first narrative form engineered natively for a fragmented attentional landscape rather than mourning it.
4. Generative models crossed the production threshold. Between 2024 and 2026, text-to-video and image-to-video systems moved from novelty to viable production tooling. The leading operators now build on video-generation engines capable of sustaining character consistency across shots — the single technical barrier that had kept synthetic drama out of commercial pipelines.
Each force alone would have produced a curiosity. Together they produced a category.
Let me be ruthless about this, because the sector's promotional literature is not.
What breaks:
What survives and becomes more valuable, not less:
There is a phrase I keep returning to, from Anne Chan of AR Productions, one of the sector's earliest operators, whose analysis of the economics is titled around a single warning: the Michelin approach will fail. It is the most useful sentence in the industry. Prestige logic — spend more, polish longer, win the review, is precisely inverted here. This is a business of velocity, iteration and portfolio, in which the quality bar is set by retention, not by craft consensus. Operators who import prestige instincts into a portfolio business will lose money with impeccable taste.
We have established what the format is, where it came from, how large it has become and why it arrived exactly now.
What we have not yet opened is the machine room, the part almost nobody outside the studios sees. How does a team of nine people produce twenty-two episodes of dark fantasy in eight weeks? How does a generative model keep a face the same across eight hundred shots? What does a minute of synthetic screen time actually cost in GPU hours, and where does that cost go?
Part II — The Machine Room: AI Production, Pipelines and Unit Economics takes the pipeline apart, layer by layer, from LoRA training to temporal deflickering, and reconstructs the true cost of a synthetic minute.

Dinis Guarda is an author, entrepreneur, founder CEO of ztudium, Businessabc, citiesabc.com and Wisdomia.ai. Dinis is an AI leader, researcher and creator who has been building proprietary solutions based on technologies like digital twins, 3D, spatial computing, AR/VR/MR. Dinis is also an author of multiple books, including "4IR AI Blockchain Fintech IoT Reinventing a Nation" and others. Dinis has been collaborating with the likes of UN / UNITAR, UNESCO, European Space Agency, IBM, Siemens, Mastercard, and governments like USAID, and Malaysia Government to mention a few. He has been a guest lecturer at business schools such as Copenhagen Business School. Dinis is ranked as one of the most influential people and thought leaders in Thinkers360 / Rise Global’s The Artificial Intelligence Power 100, Top 10 Thought leaders in AI, smart cities, metaverse, blockchain, fintech.

Beeple’s ‘Diffuse Control’ Brings AI-Powered Art to MMoCA

Blenheim Palace Brings Megalosaurus to Life Through Art, Science and Engineering

Architecture of Movement: What Deserts, Nomads and World Heritage Teach Us About Belonging

The Spiritual Meaning of the AI Singularity: What Remains Human in the Age of Artificial Intelligence?