VPU-NATIVE INFRASTRUCTURE

The era of custom silicon hardware is here.

This is not a product argument.
It is a physics argument.

The dominant pressures facing video operations today are no longer about what a system can do. They are about what it costs to keep doing it. Concurrency scales faster than audiences. Adaptive bitrate ladders expand faster than encoding efficiency improves. Power, density and predictability now matter as much as raw throughput.

VPUs dedicate 100% of silicon to video encoding.

CPU

Runs everything
as general compute.
Encodes video in software.

None.

Dedicated video-encode silicon

CPU die diagram showing no dedicated video encode silicon, illustrating that a CPU encodes video entirely in software.

=9.4 W

Bar graphic showing CPU power consumption of approximately 9.4 watts per 1080p30 video stream, the highest of the three processor types.

CPU – Power per stream

Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.

GPU

Built for graphics and Al.
A small fixed block partitioned for video encoding.

Small block.

Dedicated video-encode silicon

GPU die diagram with a small block of dedicated video encode silicon highlighted in the corner, showing that only a fixed portion of the chip handles video encoding.

=3 W

Bar graphic showing GPU power consumption of approximately 3 watts per 1080p30 video stream.

GPU – Power per stream

Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.

VPU

Built only for video.
Encode, decode, scale in video silicon.

Entire chip.

Dedicated video-encode silicon

VPU die diagram with the entire chip highlighted as dedicated video encode silicon, showing that a NETINT VPU is purpose built for video encode, decode, and scaling.

=1.56 W

Bar graphic showing NETINT VPU power consumption of approximately 1.56 watts per 1080p30 video stream, delivering up to 6x more streams per watt than a CPU.

VPU – Power per stream

Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.

Up to 6x more streams per watt.

Video infrastructure is no longer about capability. It’s now measured as cost per stream, watts per stream and streams per rack. VPUs deliver the lowest cost infrastructure.

VPUs close the codec gap.

CPU performance gains have flat lined at 5-10% each year. Meanwhile, codec complexity has increased by two orders of magnitude. The math no longer works. Purpose-built silicon is the only path to advanced codecs like AV1, at scale.

FAQ's

For Platform Engineers

What is a VPU?

A VPU (Video Processing Unit) is a chip that dedicates 100% of its silicon to video encoding. A CPU dedicates 0%. A GPU dedicates roughly 15%, with the rest built for graphics and AI. That allocation is why a VPU encodes more streams per watt than either.

Yes. NETINT VPUs run through standard FFmpeg and GStreamer interfaces, plus the libxcoder API for direct integration. There is no proprietary SDK lock-in. Your packaging and delivery stack stays in place. The encode stage moves to dedicated silicon. That is the change.

NETINT VPUs encode AV1, HEVC, and H.264 in hardware, in 8-bit and 10-bit, at resolutions up to 8K. They also decode VP9. AV1 is the one that matters most: it carries 100 to 200x the compute complexity of H.264, which is past the point where software encoding scales economically.

Compare it at the presets you actually run in production, not benchmark presets nobody ships. At production speed settings and matched bitrates, VPU output is comparable. Run the comparison on your own content during a pilot. We would rather you verify than take our word.

Start in the cloud at no cost. Ecosystem partner clouds fund VPU pilots: Akamai Cloud applies $500 in credit to a 30-day pilot for new accounts, and NetActuate runs no-cost proofs of concept for qualified teams. You qualify the architecture against your own workloads before any hardware decision.

For Operations Leaders

How many streams can a VPU server handle?

A NETINT Quadra Video Server encodes 320 simultaneous 1080p30 streams, 80 at 4Kp30, or 20 at 8Kp30. That is roughly 40x the output of a CPU-based encoding server in the same rack space.

Every VPU has a fixed, known encoding capacity. The card that handles 320 streams in hour one handles 320 streams in hour 10,000. You plan procurement, facility expansion, and staffing on counts, not on elastic headroom estimates.

No. VPUs are PCIe and U.2 cards that install in servers you already own, including chassis you have decommissioned. You can also run VPU instances on partner clouds, or combine both. The architecture is deployment-agnostic.

The migration runs in parallel. VPU encoding runs alongside your current encoding on the same content while you compare output and cost. Cutover happens only when your own data says so, and every step is reversible. Explore takes 1 to 2 weeks, testing 2 to 4, migration 4 to 8.

Measured at the card, a Quadra T1A draws 20 watts at full load while encoding 32 simultaneous 1080p30 streams. That is 0.63 watts per stream for the VPU itself.

Measured at the wall, a fully loaded 1RU Quadra Video Server draws about 500 watts while encoding 320 streams. That is about 1.56 watts per stream for the complete system. Every figure here is measured, and you can see how on our /Benchmarks page.

For Finance & Strategy

How much does VPU encoding reduce cost?

Per-stream encoding cost drops 60 to 80% versus software encoding on general-purpose servers. The savings compound with scale: cost per stream stays flat as concurrency grows, where general-purpose cost climbs.

Three curves crossed. Codec complexity grew 100x with AV1. CPU performance grows 5 to 10% a year. And energy regulation, especially in Europe, now puts a compliance deadline on watts per stream. Waiting does not pause any of them.

Either. Owned VPU cards in your own or existing chassis are capex. VPU instances on partner clouds are opex. Most converging companies run hybrid. The economics work in all three models because the efficiency lives in the silicon, not the deployment.

Google and Meta built their own encoding ASICs (Argos, MSVP) over decades of capital investment, and none of it is for sale. VPUs are the open market’s route to the same architecture: purpose-built silicon, available to every platform that competes openly.

One to two weeks. The pilot runs your real workloads and produces your numbers: cost per stream, watts per stream, streams per rack. The investment case is built from your data, not our claims.

vpu ecosystem

VPU-native infrastructure is the model. The VPU Ecosystem makes it deployable.

Ready to go VPU-native?

Moving to VPU-native infrastructure is not a technology decision. 
It is an infrastructure decision requiring internal alignment.
The Deployment Playbook is how converging companies got there.