Building a Category in Plain Sight: A Decade Inside the Rise of the Video Processing Unit

On the evening of October 9, 2024, our team gathered in the Prince George Ballroom in New York for the 75th Technology and Engineering Emmy Awards. The category was new to most people in our industry: “Design and Deployment of Efficient Hardware Video Accelerators for Cloud.” There were four recipients on the list. Three of them were AMD, Google, and Meta. The fourth was NETINT.

I want to be clear about what that meant and what it did not. It did not mean we had outbuilt the largest engineering organizations in the world. It meant that the work we had been doing since 2015, alongside the parallel work those companies had been doing inside their own walls, had quietly added up to a new category of silicon, the VPU or video processing unit. The Television Academy’s recognition was not a verdict on any one chip. It was an acknowledgment that the way the world processes video had changed, and that purpose-built hardware now sits at the center of it.

That is the story I want to tell. Not because NETINT needs more credit, but because the path that got us here is about restraint as much as it is about silicon.

A short history of the silicon under your screen

Most of the video that reaches a phone, a television, or a laptop has been encoded multiple times before it arrives. That work used to happen on general-purpose computers. For the better part of two decades, the open-source x264 encoder set the standard for software encoding on CPUs, and it remains in wide use today. The economics worked when video was a small fraction of internet traffic and audiences were measured in millions rather than billions.

The next chapter belonged to GPUs. NVIDIA introduced its NVENC dedicated encoder block with the Kepler architecture, and successive generations added support for HEVC, AV1, and higher resolutions. NVENC is technically separate from the CUDA cores on the chip, which is to say even GPU vendors recognized that video encoding deserved its own silicon. AMD and Intel followed with their own integrated encoders. These solutions made certain workloads viable that had been impossible on CPUs alone.

A parallel track ran through field-programmable gate arrays, or FPGAs. Each of these eras solved a real problem. None of them solved the problem that emerged when video did not just become a major share of internet traffic but became the substrate of the internet itself.

Why a fourth category had to exist

By 2021, YouTube was uploading more than 700 hours of video every minute, and platforms built on user-generated short-form content, cloud gaming, virtual desktops, and personalized ad insertion were piling on demand that grew faster than Moore’s law could absorb. (By that same year, more than 200 billion minutes of video had been processed by NETINT silicon, a figure the analyst Dylan Patel published in his 2022 SemiAnalysis profile of NETINT.)

Inside Google, this challenge was so acute that the company began a custom silicon program in 2015. Six years later, the world saw the result: Argos, a Video Coding Unit deployed across YouTube infrastructure that, by Google’s own published comparisons, was 20 to 33 times more compute-efficient than its previous server fleet for the same encoding workloads. Meta’s first internal video ASIC, Mount Shasta, followed in 2019, and the company later disclosed its second-generation Meta Scalable Video Processor, which Meta engineers said delivered roughly nine times the H.264 throughput of libx264 software encoding while drawing about ten watts of power.

The hyperscalers had reached the same conclusion that our co-founders, Alex Liu and Tao Zhong, had reached in 2015. Video encoding was no longer a workload that should ride on top of general-purpose silicon. It was a workload that deserved its own vehicle. The difference was that the largest platforms were building for themselves. We were building for everyone else.

The pivot that built the company

NETINT did not start as a video silicon company. Our co-founders set out in 2015 to design a new generation of solid-state drives, and the early work culminated in what was the world’s first PCIe Gen 4 computational storage device. The market for that product turned out to be smaller and more concentrated than the team had hoped. The asset we had built, however, was a deep capability in custom ASIC design and the unusual idea of putting compute capability on a storage form factor.

In 2017, we began converting that capability into video silicon. In 2018, we shipped the T408, the first commercially available ASIC-based video transcoder. It drew about seven watts under full load, occupied a U.2 NVMe slot, and gave a single server the ability to transcode dozens of simultaneous high-definition streams in real time.

We tell this story carefully because the lesson inside it matters more than the product. The pivot was not from one market to another. It was from a category we wanted to be in to a category that wanted us. The NVMe form factor we had designed for storage turned out to be the right physical envelope for high-density video silicon, because it slotted directly into existing servers and existing data center workflows without forcing customers to rebuild their racks. That accident of architecture became one of our most durable advantages.

What category creation actually looks like

Most discussions of category creation focus on naming and positioning. Those activities matter, but they are downstream of two harder choices. The first is whether the category genuinely exists, or whether you are inventing language for a problem that customers do not feel. The second is whether you are willing to be patient enough for the world to catch up. Put another way, are you willing to put in the evangelistic effort to teach the market that what they are doing must be changed.

We made the bet that the category genuinely existed. The largest platforms were quietly reaching the same conclusion in their own labs. Cisco was projecting that video would account for the majority of internet traffic. Energy costs were tightening data center planning, and the economics of running general-purpose CPUs against codecs like HEVC and AV1 were showing strain. We did not need to create demand. We needed to build the merchant version of a capability that the largest platforms in the world would soon need to license, build, or rent.

The second bet was harder. The first generation of our chip, codenamed Logan and built on TSMC 28 nanometers, was a product. The second generation, Quadra, built on Samsung 14 nanometers, was the foundation of the category we now talk about. It quadrupled encoding capacity, added AI inference and 2D engines, and was the first hardware implementation of AV1 encoding in silicon, which we shipped commercially in early 2021. Customers do not wake up ready to adopt a new compute category. We spent years writing reference architectures, publishing benchmarks, and integrating into FFmpeg, GStreamer, and the toolchains video engineers already trusted. Patience here is not optional. It is the strategy.

What customers actually do with a VPU

I am cautious about customer claims because the most credible thing a vendor can do is point to outside evidence. In his 2022 SemiAnalysis profile, Dylan Patel wrote that NETINT had “racked up a number of major wins including major names such as ByteDance, Baidu, Tencent, Alibaba, Kuaishou, and a similar sized US based global platform.” That paragraph remains one of the most-cited descriptions of our customer base, and we did not write it.

In broadcast, Easy Tools, the Paris-based modular video processing company replaced ten CPU-based servers with a single Quadra video server in their broadcast encoding workflow. In a published interview on our Voices of Video podcast, CEO Christophe Massiot described the form factor decision in clear terms: “NETINT made a very clever choice: the NVMe form factor means on a regular one-rack-unit server you can put eight or ten modules, so that’s a game-changer.”

In cloud delivery, our partnership with Akamai, announced in September 2024, brings VPU-accelerated compute to one of the most distributed cloud networks in the world. Published benchmarks from that integration, conducted with the independent test lab Cires21, showed energy efficiency improvements of four to six times against GPU alternatives, with the VPU consuming roughly 0.4 to 0.7 watts per stream where comparable GPUs consumed 2.5 to 3.6 watts per stream. Quality, measured by the VMAF perceptual metric, was held constant.

These are not isolated proof points. They reflect the same pattern across categories of customer, from short-form social platforms to broadcasters to live event distributors and cloud gaming platforms. When a workload is dominated by one task and that task is encoding video, putting it on silicon designed for that task produces a step change in cost and energy that procurement discounts alone cannot match.

The four cost levers we hear most about

When customers describe what is breaking, they tend to describe four things at once. Compute scales linearly with concurrency, so doubling the audience doubles the fleet. Power and density turn into hard ceilings inside data centers and into pricing volatility inside cloud regions. Egress fees on transcoded video often exceed the cost of the transcoding itself. And the operational complexity of managing thousands of general-purpose instances quietly absorbs engineering time that should be spent on product.

A VPU does not solve all four. It addresses the first two directly, by encoding more streams per watt and per rack unit than software running on CPUs or GPUs. The third, egress, only resolves when the compute layer sits inside a delivery network designed for video economics. The fourth, operational complexity, only resolves when the silicon is wrapped in a software stack engineers can adopt without rebuilding their pipelines. That is why we have invested heavily in our Bitstreams platform, FFmpeg and GStreamer integrations, and our partnership with Akamai as we have in the chip itself.

What the Emmy actually meant

Industry awards are easy to overstate, but in 2024, what the Television Academy was actually acknowledging on behalf of the broadcast and streaming community is that the move to purpose-built video silicon was consequential enough to be recognized. Sharing that honor with AMD, Google, and Meta placed the conversation about VPUs in the same frame as the most significant infrastructure investments in the industry. After years of debunking hardware myths from the industry, this was a significant recognition of the VPU category.

It also clarified something for our customers and partners. For years, the question we were asked most often was whether VPUs were a viable category, or a wedge product that GPUs and CPUs would eventually absorb. The shared award answered that question in a way no marketing campaign could. The largest video platforms in the world had built their own. We had built one anyone could buy. Both paths now exist alongside each other, and both validate the underlying thesis. Video encoding is moving from software to hardware for all but the most niche applications.

The lesson for founders

If there is a playbook embedded in this story, it is shorter than people expect.

Build for a category the world will eventually need, not the one it currently asks for. The largest platforms’ internal silicon programs were running in parallel with ours, sometimes without our knowledge. We did not need to convince them the problem was real. We needed to be ready when the rest of the market arrived at the same understanding.

Choose a form factor that meets the world where it already is. Our decision to build a video accelerator that fits in an NVMe slot was not glamorous. But it is one of the reasons we are deployed in production at scale today. As of this writing, there are more than 200,000 NETINT VPUs in product. One well known competitor that did not choose this path was forced to exit the category due to the tremendously difficult integration hurdles they faced.

Be honest about what your technology does not solve. CPUs and GPUs still have a role in many video workflows. The right framing for a VPU is not a replacement of everything else. It is acceleration of a specific, dominant workload, with the discipline to integrate cleanly with what surrounds it.

Invest in the ecosystem before the ecosystem returns the favor. Our work inside the Alliance for Open Media, our published benchmarks, our podcast and content programs, and our partnerships with Akamai, and members of the VPU Ecosystem were not separate from the silicon strategy. They were the silicon strategy.

Above all, be patient. The first commercial ASIC video transcoder shipped in 2018. The first commercial AV1 hardware shipped in 2021. The Tech Emmy came in 2024. None of those steps happened on a quarterly cadence, and none of them would have happened if we had treated them as marketing milestones rather than category defining infrastructure work.

Where this goes next

The economics of video are still tightening. AV1 adoption is broadening, AI-driven workflows are adding new compute demands at the encoder and decoder, and the geographic expansion of streaming is making density and energy more, not less, important. The future of this category is not just denser silicon. It is silicon that is more deeply integrated into distributed compute, more closely coupled to AI inference, and more directly accountable to the unit economics that decide whether a video business can scale profitably.

I am not going to claim we have all of that figured out. What I will claim is that the category is real, the recognition is shared, and the work continues. If you are an operator trying to decide whether the move to purpose-built video silicon is worth the effort, the answer is in your own infrastructure bills and energy reports more than it is in any vendor’s deck. We are happy to talk about what that math looks like for your workload, and we will be straightforward about where a VPU helps and where it does not.

That is the playbook. Find a category the future is already building toward. Build the merchant version. Be patient enough to let the market catch up. And when the recognition finally comes, share the stage gracefully, because the people you are sharing it with usually got there for the same reasons you did.

ACCESS NOW:  ASIC-Based Transcoding
for High-volume Use Cases
Including social media, broadcast, interactive platforms, and service providers


ACCESS NOW

Building a Category in Plain Sight: A Decade Inside the Rise of the Video Processing Unit

A decade inside the rise of the video processing unit, from industry skepticism to Emmy recognition, and how the VPU reshaped streaming infrastructure.

On the evening of October 9, 2024, our team gathered in the Prince George Ballroom in New York for the 75th Technology and Engineering Emmy Awards. The category was new to most people in our industry: “Design and Deployment of Efficient Hardware Video Accelerators for Cloud.” There were four recipients on the list. Three of them were AMD, Google, and Meta. The fourth was NETINT.

I want to be clear about what that meant and what it did not. It did not mean we had outbuilt the largest engineering organizations in the world. It meant that the work we had been doing since 2015, alongside the parallel work those companies had been doing inside their own walls, had quietly added up to a new category of silicon, the VPU or video processing unit. The Television Academy’s recognition was not a verdict on any one chip. It was an acknowledgment that the way the world processes video had changed, and that purpose-built hardware now sits at the center of it.

That is the story I want to tell. Not because NETINT needs more credit, but because the path that got us here is about restraint as much as it is about silicon.

A short history of the silicon under your screen

Most of the video that reaches a phone, a television, or a laptop has been encoded multiple times before it arrives. That work used to happen on general-purpose computers. For the better part of two decades, the open-source x264 encoder set the standard for software encoding on CPUs, and it remains in wide use today. The economics worked when video was a small fraction of internet traffic and audiences were measured in millions rather than billions.

The next chapter belonged to GPUs. NVIDIA introduced its NVENC dedicated encoder block with the Kepler architecture, and successive generations added support for HEVC, AV1, and higher resolutions. NVENC is technically separate from the CUDA cores on the chip, which is to say even GPU vendors recognized that video encoding deserved its own silicon. AMD and Intel followed with their own integrated encoders. These solutions made certain workloads viable that had been impossible on CPUs alone.

A parallel track ran through field-programmable gate arrays, or FPGAs. Each of these eras solved a real problem. None of them solved the problem that emerged when video did not just become a major share of internet traffic but became the substrate of the internet itself.

Why a fourth category had to exist

By 2021, YouTube was uploading more than 700 hours of video every minute, and platforms built on user-generated short-form content, cloud gaming, virtual desktops, and personalized ad insertion were piling on demand that grew faster than Moore’s law could absorb. (By that same year, more than 200 billion minutes of video had been processed by NETINT silicon, a figure the analyst Dylan Patel published in his 2022 SemiAnalysis profile of NETINT.)

Inside Google, this challenge was so acute that the company began a custom silicon program in 2015. Six years later, the world saw the result: Argos, a Video Coding Unit deployed across YouTube infrastructure that, by Google’s own published comparisons, was 20 to 33 times more compute-efficient than its previous server fleet for the same encoding workloads. Meta’s first internal video ASIC, Mount Shasta, followed in 2019, and the company later disclosed its second-generation Meta Scalable Video Processor, which Meta engineers said delivered roughly nine times the H.264 throughput of libx264 software encoding while drawing about ten watts of power.

The hyperscalers had reached the same conclusion that our co-founders, Alex Liu and Tao Zhong, had reached in 2015. Video encoding was no longer a workload that should ride on top of general-purpose silicon. It was a workload that deserved its own vehicle. The difference was that the largest platforms were building for themselves. We were building for everyone else.

The pivot that built the company

NETINT did not start as a video silicon company. Our co-founders set out in 2015 to design a new generation of solid-state drives, and the early work culminated in what was the world’s first PCIe Gen 4 computational storage device. The market for that product turned out to be smaller and more concentrated than the team had hoped. The asset we had built, however, was a deep capability in custom ASIC design and the unusual idea of putting compute capability on a storage form factor.

In 2017, we began converting that capability into video silicon. In 2018, we shipped the T408, the first commercially available ASIC-based video transcoder. It drew about seven watts under full load, occupied a U.2 NVMe slot, and gave a single server the ability to transcode dozens of simultaneous high-definition streams in real time.

We tell this story carefully because the lesson inside it matters more than the product. The pivot was not from one market to another. It was from a category we wanted to be in to a category that wanted us. The NVMe form factor we had designed for storage turned out to be the right physical envelope for high-density video silicon, because it slotted directly into existing servers and existing data center workflows without forcing customers to rebuild their racks. That accident of architecture became one of our most durable advantages.

What category creation actually looks like

Most discussions of category creation focus on naming and positioning. Those activities matter, but they are downstream of two harder choices. The first is whether the category genuinely exists, or whether you are inventing language for a problem that customers do not feel. The second is whether you are willing to be patient enough for the world to catch up. Put another way, are you willing to put in the evangelistic effort to teach the market that what they are doing must be changed.

We made the bet that the category genuinely existed. The largest platforms were quietly reaching the same conclusion in their own labs. Cisco was projecting that video would account for the majority of internet traffic. Energy costs were tightening data center planning, and the economics of running general-purpose CPUs against codecs like HEVC and AV1 were showing strain. We did not need to create demand. We needed to build the merchant version of a capability that the largest platforms in the world would soon need to license, build, or rent.

The second bet was harder. The first generation of our chip, codenamed Logan and built on TSMC 28 nanometers, was a product. The second generation, Quadra, built on Samsung 14 nanometers, was the foundation of the category we now talk about. It quadrupled encoding capacity, added AI inference and 2D engines, and was the first hardware implementation of AV1 encoding in silicon, which we shipped commercially in early 2021. Customers do not wake up ready to adopt a new compute category. We spent years writing reference architectures, publishing benchmarks, and integrating into FFmpeg, GStreamer, and the toolchains video engineers already trusted. Patience here is not optional. It is the strategy.

What customers actually do with a VPU

I am cautious about customer claims because the most credible thing a vendor can do is point to outside evidence. In his 2022 SemiAnalysis profile, Dylan Patel wrote that NETINT had “racked up a number of major wins including major names such as ByteDance, Baidu, Tencent, Alibaba, Kuaishou, and a similar sized US based global platform.” That paragraph remains one of the most-cited descriptions of our customer base, and we did not write it.

In broadcast, Easy Tools, the Paris-based modular video processing company replaced ten CPU-based servers with a single Quadra video server in their broadcast encoding workflow. In a published interview on our Voices of Video podcast, CEO Christophe Massiot described the form factor decision in clear terms: “NETINT made a very clever choice: the NVMe form factor means on a regular one-rack-unit server you can put eight or ten modules, so that’s a game-changer.”

In cloud delivery, our partnership with Akamai, announced in September 2024, brings VPU-accelerated compute to one of the most distributed cloud networks in the world. Published benchmarks from that integration, conducted with the independent test lab Cires21, showed energy efficiency improvements of four to six times against GPU alternatives, with the VPU consuming roughly 0.4 to 0.7 watts per stream where comparable GPUs consumed 2.5 to 3.6 watts per stream. Quality, measured by the VMAF perceptual metric, was held constant.

These are not isolated proof points. They reflect the same pattern across categories of customer, from short-form social platforms to broadcasters to live event distributors and cloud gaming platforms. When a workload is dominated by one task and that task is encoding video, putting it on silicon designed for that task produces a step change in cost and energy that procurement discounts alone cannot match.

The four cost levers we hear most about

When customers describe what is breaking, they tend to describe four things at once. Compute scales linearly with concurrency, so doubling the audience doubles the fleet. Power and density turn into hard ceilings inside data centers and into pricing volatility inside cloud regions. Egress fees on transcoded video often exceed the cost of the transcoding itself. And the operational complexity of managing thousands of general-purpose instances quietly absorbs engineering time that should be spent on product.

A VPU does not solve all four. It addresses the first two directly, by encoding more streams per watt and per rack unit than software running on CPUs or GPUs. The third, egress, only resolves when the compute layer sits inside a delivery network designed for video economics. The fourth, operational complexity, only resolves when the silicon is wrapped in a software stack engineers can adopt without rebuilding their pipelines. That is why we have invested heavily in our Bitstreams platform, FFmpeg and GStreamer integrations, and our partnership with Akamai as we have in the chip itself.

What the Emmy actually meant

Industry awards are easy to overstate, but in 2024, what the Television Academy was actually acknowledging on behalf of the broadcast and streaming community is that the move to purpose-built video silicon was consequential enough to be recognized. Sharing that honor with AMD, Google, and Meta placed the conversation about VPUs in the same frame as the most significant infrastructure investments in the industry. After years of debunking hardware myths from the industry, this was a significant recognition of the VPU category.

It also clarified something for our customers and partners. For years, the question we were asked most often was whether VPUs were a viable category, or a wedge product that GPUs and CPUs would eventually absorb. The shared award answered that question in a way no marketing campaign could. The largest video platforms in the world had built their own. We had built one anyone could buy. Both paths now exist alongside each other, and both validate the underlying thesis. Video encoding is moving from software to hardware for all but the most niche applications.

The lesson for founders

If there is a playbook embedded in this story, it is shorter than people expect.

Build for a category the world will eventually need, not the one it currently asks for. The largest platforms’ internal silicon programs were running in parallel with ours, sometimes without our knowledge. We did not need to convince them the problem was real. We needed to be ready when the rest of the market arrived at the same understanding.

Choose a form factor that meets the world where it already is. Our decision to build a video accelerator that fits in an NVMe slot was not glamorous. But it is one of the reasons we are deployed in production at scale today. As of this writing, there are more than 200,000 NETINT VPUs in product. One well known competitor that did not choose this path was forced to exit the category due to the tremendously difficult integration hurdles they faced.

Be honest about what your technology does not solve. CPUs and GPUs still have a role in many video workflows. The right framing for a VPU is not a replacement of everything else. It is acceleration of a specific, dominant workload, with the discipline to integrate cleanly with what surrounds it.

Invest in the ecosystem before the ecosystem returns the favor. Our work inside the Alliance for Open Media, our published benchmarks, our podcast and content programs, and our partnerships with Akamai, and members of the VPU Ecosystem were not separate from the silicon strategy. They were the silicon strategy.

Above all, be patient. The first commercial ASIC video transcoder shipped in 2018. The first commercial AV1 hardware shipped in 2021. The Tech Emmy came in 2024. None of those steps happened on a quarterly cadence, and none of them would have happened if we had treated them as marketing milestones rather than category defining infrastructure work.

Where this goes next

The economics of video are still tightening. AV1 adoption is broadening, AI-driven workflows are adding new compute demands at the encoder and decoder, and the geographic expansion of streaming is making density and energy more, not less, important. The future of this category is not just denser silicon. It is silicon that is more deeply integrated into distributed compute, more closely coupled to AI inference, and more directly accountable to the unit economics that decide whether a video business can scale profitably.

I am not going to claim we have all of that figured out. What I will claim is that the category is real, the recognition is shared, and the work continues. If you are an operator trying to decide whether the move to purpose-built video silicon is worth the effort, the answer is in your own infrastructure bills and energy reports more than it is in any vendor’s deck. We are happy to talk about what that math looks like for your workload, and we will be straightforward about where a VPU helps and where it does not.

That is the playbook. Find a category the future is already building toward. Build the merchant version. Be patient enough to let the market catch up. And when the recognition finally comes, share the stage gracefully, because the people you are sharing it with usually got there for the same reasons you did.

ACCESS NOW:  ASIC-Based Transcoding
for High-volume Use Cases
Including social media, broadcast, interactive platforms, and service providers


ACCESS NOW