This website uses cookies and other technologies to help us provide you with better content and customized services. If you want to continue to enjoy this website’s content, please agree to our use of cookies. For more information on cookies and their use, please see our latest Privacy Policy.

Accept

cwlogo

切換側邊選單 切換搜尋選單

Meta's Message at SEMICON Taiwan 2026: Taiwan's Semiconductor Supply Chain Essential for the AI Revolution

Meta's Message at SEMICON Taiwan 2026: Taiwan's Semiconductor Supply Chain Essential for the AI Revolution

Source:SEMI

When a World Cup match turned into a stunning comeback and 30 million messages per second swarmed across the globe, how did Meta's systems quietly absorb it all? At the SEMICON Taiwan 2026 CEO Summit & Ecosystem Executive Summit, CQ (Chunqiang) Tang, VP of Engineering, AI & Compute Foundation at Meta, provided a behind-the-scenes look at Meta's in-house chip development, the relocation of entire data centers, and the machinery behind gigawatt-scale AI compute.

Views

99+
Share

Meta's Message at SEMICON Taiwan 2026: Taiwan's Semiconductor Supply Chain Essential for the AI Revolution

By SEMI
Sponsored Content

CQ (Chunqiang) Tang, VP of Engineering, AI & Compute Foundation at Meta, speaking at the SEMICON Taiwan 2026 CEO Summit & Ecosystem Executive Summit about how Meta uses hardware-software co-design and automated orchestration to deliver resilient, elastic AI infrastructure.

In the round of 16 at the 2026 World Cup, Argentina and Egypt faced off. Egypt came out firing and went 2–0 up, but Argentina showed remarkable persistence, scoring three goals in the final 15 minutes for a 3–2 comeback win. The stunning turnaround set off a wave of conversation among fans worldwide, driving WhatsApp traffic to nearly 30 million messages per second.

"That number is larger than the entire population of Taiwan," noted CQ (Chunqiang) Tang, VP of Engineering, AI & Compute Foundation at Meta, in his keynote at SEMICON Taiwan 2026's CEO Summit & Ecosystem Executive Summit. The smooth handling of that sudden spike in online activity, he said, is a snapshot of the resilience built into Meta's underlying infrastructure.

30 million messages a second: Meta's auto-scaling quietly tames a huge spike in traffic

The ability to seamlessly absorb the sudden surge, without any human intervention, comes down to the sheer scale of Meta's global services and its automated orchestration.

Meta's family of apps serves up to 3.6 billion people a day—more than 40% of the world's population—with Reels plays exceeding 200 billion a day and WhatsApp delivering more than 100 billion messages every day, supported by dozens of regional data centers across North America, Europe, and Asia. When an unforeseen wave of traffic hits, the system triggers an auto-scaling mechanism. Within seconds, it dynamically reallocates resources, pausing lower-priority background tasks and steering compute capacity to user-facing activity to keep core communications responsive.

This elastic infrastructure does more than absorb sudden traffic spikes—it also allows Meta to deploy quickly and scale on demand. The showcase example is Threads, the social platform that has grown rapidly in recent years.

"The product team built Threads in just five months, but the infrastructure team was officially told about the launch only two days before it went live," CQ said with a laugh. Most companies would struggle to even draw up a plan in two days, let alone execute the launch of a new platform, he observed, before adding, "But we actually did it."

CQ emphasized Taiwan's central role in the semiconductor supply chain and Meta's exploration of a range of possibilities for deeper collaboration.

The war-room culture behind the fastest-ever ramp to 100 million users

Drawing on both company culture and infrastructure strengths, Threads reached 100 million users in five days—a world record. By comparison, Twitter (formerly Twitter) took five years, TikTok nine months, and even ChatGPT two months. Threads has now passed 500 million monthly active users, placing it among the largest text-based social platforms—and one that's especially popular in Taiwan.

Behind all this is more than a decade of accumulated experience in auto-scaling and automated disaster recovery—and, just as critically, Meta's distinctive war-room culture. When a new product launches at a scale of hundreds of millions of users, failures are inevitable, no matter how good the automation technology. Round-the-clock war rooms staffed by global teams, combined with Meta's "move fast" mindset, are what turned the seemingly impossible into reality.

Everything at Meta, from feed recommendations and short-video rankings to real-time interaction on Threads and the AI assistant in smart glasses, relies on an enormous and highly heterogeneous set of AI workloads—each with very different demands on compute, memory, and network bandwidth.

From in-house MTIA chips to cross-layer co-design

"One size doesn't fit all," CQ said. To support varied and fast-evolving AI workloads, Meta takes a flexible approach to hardware, deploying GPUs from both NVIDIA and AMD alongside MTIA, its own purpose-built AI accelerator. With Moore's Law no longer delivering the gains it once did, single-point improvements are no longer enough, so Meta applies cross-layer co-design throughout its technology stack—which includes chips, server racks, networking, software, and AI models—to maximize performance, reliability, and energy efficiency.

Meta has in fact been designing its own chips for many years. In 2023, its work on high-performance video processing silicon was recognized with a Technology & Engineering Emmy Award for Design and Deployment of Efficient Hardware Video Accelerators for Cloud.

The company has developed multiple generations of MTIA to date and has hundreds of thousands of units deployed in production. But the biggest advantage of in-house silicon, CQ stressed, is that it anchors full-stack co-design—from chips to servers and from software to AI models. Vertical integration enables Meta to extract maximum performance from its hardware and software by tuning them comprehensively for the actual workloads created by billions of users.

This matters most in AI inference: for every token a model generates, it must draw on its trained knowledge and consider the entire context so far. The volume of data that must be moved makes memory bandwidth the decisive factor in how quickly a user gets a response—and this is why breaking through the memory wall is driving advances in advanced packaging and optical interconnects.

Building gigawatt-scale AI compute

In-house silicon alone is not enough, however. Training and running powerful AI models has pushed compute demand well beyond previous physical space constraints, from individual chips all the way to hyperscale clusters spanning multiple data centers.

In 2024, to assemble a cluster of more than 100,000 GPUs, Meta evacuated  five operating data centers and relocated their more than 5,000 server racks without affecting the user experience—even developing new robots to help with the job. It was able to quadruple network bandwidth, ultimately standing up a 120,000-GPU cluster within months.

And that was only the starting point for Meta's compute expansion. After the 120,000-GPU cluster was complete, CQ said, the team kept pushing the limits of what could be built.

Meta expects to bring Prometheus, a cluster with more than 1 GW of power capacity, online by the end of this year, and over the next few years plans to do the same for Hyperion, a cluster with a power capacity of up to 5 GW—formally taking AI infrastructure into the gigawatt era.

Tackling silent data corruption, a stealthy threat to AI GPU clusters

But at gigawatt scale, the hardest test of all is keeping operations reliable.

CQ likened tens of thousands of GPUs training in lockstep to 10,000 race cars chained together, all speeding down the track as a single mass: if one should fail, all need to stop. Working with partners over the past year, Meta has cut failure rates by a factor of 50.

Beyond outright hardware failures, silent data corruption (SDC) is a threat that's far more difficult to detect. A chip can generate the wrong results without triggering a hardware alert at all, which is devastating for a training run. Meta has therefore developed a range of technologies, including the open-source CP-Bench tool, to detect GPU faults in real time and isolate them to prevent a system-wide outage.

Partnering with Taiwan's semiconductor companies for the AI revolution

Getting past these frontier technology bottlenecks is not something Meta can do alone. "Taiwan sits at the core of the global semiconductor supply chain, and the success of Meta's AI strategy depends heavily on the support of Taiwanese companies," said CQ, who was representing Meta at SEMICON Taiwan for the first time. What makes the event unique, he added, is the breadth and depth of its ecosystem participants.

"We can not only deepen our engagement with existing partners such as NVIDIA and Broadcom here, but also step outside our usual circles and connect directly with many more equipment manufacturers, materials suppliers, and other potential partners across the technology supply chain," CQ said.

To address the next generation of challenges—such as memory wall, advanced packaging, and optical interconnects—SEMICON Taiwan brings together key participants from industry, government, academia, and research organizations on a remarkable scale. CQ praised the organizers' professional execution of the event and emphasized that Meta will continue to work through open standards and open-source tools such as the Open Compute Project, PyTorch, and its latest open-weight model, Muse Glimmer, to deepen collaboration with Taiwan's semiconductor supply chain and build a strong, open AI future together.


FAQ

Q1: How does Meta keep its services available in the face of sudden, massive traffic spikes?

A: Meta uses auto-scaling and dynamic resource allocation. Within seconds, these mechanisms kick in, pausing lower-priority background jobs and steering compute capacity to user-facing activity to keep core communications responsive.

Q2: What is Meta's chip strategy for such a diverse array of AI workloads?

A: Meta takes a flexible approach, using GPUs from both NVIDIA and AMD alongside MTIA, its own purpose-built AI accelerator. Meta also applies cross-layer co-design, optimizing its full stack of hardware and software technologies—from chips to servers and from networks to AI models.

Q3: How does Meta handle silent data corruption (SDC) in gigawatt-scale AI clusters?

A: SDC causes a chip to silently produce incorrect results and can destroy an entire training run; it is a challenge the whole industry faces with hyperscale GPU clusters. Meta has therefore developed a range of technologies—including the open-source CP-Bench tool—to detect GPU faults in real time and isolate them to prevent a system-wide outage.

Q4: Why did Meta emphasize at SEMICON Taiwan 2026 that its AI strategy depends heavily on Taiwan's semiconductor ecosystem?

A: Meta noted that Taiwan sits at the core of the semiconductor supply chain and that SEMICON Taiwan offers uniquely broad and deep access to the ecosystem, including direct connections with equipment manufacturers, materials suppliers, and other potential partners. To address next-generation challenges such as memory wall, advanced packaging, and optical interconnects, SEMICON Taiwan, organized by SEMI, leverages its global industry network to connect technology companies with Taiwan's semiconductor supply chain. It therefore serves as an important platform for facilitating technology exchange and ecosystem collaboration. Meta used the event to emphasize that it will explore a range of collaboration possibilities through open standards and open-source tools, including the Open Compute Project, PyTorch, and its latest open-weight model, Muse Glimmer.

Views

99+
Share

Keywords:

好友人數