The dominant pressures facing video operations today are no longer about what a system can do. They are about what it costs to keep doing it. Concurrency scales faster than audiences. Adaptive bitrate ladders expand faster than encoding efficiency improves. Power, density and predictability now matter as much as raw throughput.
Runs everything
as general compute.
Encodes video in software.
None.
Dedicated video-encode silicon
CPU – Power per stream
Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.
Built for graphics and Al.
A small fixed block partitioned for video encoding.
Small block.
Dedicated video-encode silicon
GPU – Power per stream
Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.
Built only for video.
Encode, decode, scale in video silicon.
Entire chip.
Dedicated video-encode silicon
VPU – Power per stream
Based on total svstem power divided by sustained 1080p30 encoding capacity. Results will vary.
Up to 6x more streams per watt.
CPU performance gains have flat lined at 5-10% each year. Meanwhile, codec complexity has increased by two orders of magnitude. The math no longer works. Purpose-built silicon is the only path to advanced codecs like AV1, at scale.
A VPU (Video Processing Unit) is a chip that dedicates 100% of its silicon to video encoding. A CPU dedicates 0%. A GPU dedicates roughly 15%, with the rest built for graphics and AI. That allocation is why a VPU encodes more streams per watt than either.
Yes. NETINT VPUs run through standard FFmpeg and GStreamer interfaces, plus the libxcoder API for direct integration. There is no proprietary SDK lock-in. Your packaging and delivery stack stays in place. The encode stage moves to dedicated silicon. That is the change.
NETINT VPUs encode AV1, HEVC, and H.264 in hardware, in 8-bit and 10-bit, at resolutions up to 8K. They also decode VP9. AV1 is the one that matters most: it carries 100 to 200x the compute complexity of H.264, which is past the point where software encoding scales economically.
Compare it at the presets you actually run in production, not benchmark presets nobody ships. At production speed settings and matched bitrates, VPU output is comparable. Run the comparison on your own content during a pilot. We would rather you verify than take our word.
Start in the cloud at no cost. Ecosystem partner clouds fund VPU pilots: Akamai Cloud applies $500 in credit to a 30-day pilot for new accounts, and NetActuate runs no-cost proofs of concept for qualified teams. You qualify the architecture against your own workloads before any hardware decision.
A NETINT Quadra Video Server encodes 320 simultaneous 1080p30 streams, 80 at 4Kp30, or 20 at 8Kp30. That is roughly 40x the output of a CPU-based encoding server in the same rack space.
Every VPU has a fixed, known encoding capacity. The card that handles 320 streams in hour one handles 320 streams in hour 10,000. You plan procurement, facility expansion, and staffing on counts, not on elastic headroom estimates.
No. VPUs are PCIe and U.2 cards that install in servers you already own, including chassis you have decommissioned. You can also run VPU instances on partner clouds, or combine both. The architecture is deployment-agnostic.
The migration runs in parallel. VPU encoding runs alongside your current encoding on the same content while you compare output and cost. Cutover happens only when your own data says so, and every step is reversible. Explore takes 1 to 2 weeks, testing 2 to 4, migration 4 to 8.
Measured at the card, a Quadra T1A draws 20 watts at full load while encoding 32 simultaneous 1080p30 streams. That is 0.63 watts per stream for the VPU itself.
Measured at the wall, a fully loaded 1RU Quadra Video Server draws about 500 watts while encoding 320 streams. That is about 1.56 watts per stream for the complete system. Every figure here is measured, and you can see how on our /Benchmarks page.
Per-stream encoding cost drops 60 to 80% versus software encoding on general-purpose servers. The savings compound with scale: cost per stream stays flat as concurrency grows, where general-purpose cost climbs.
Three curves crossed. Codec complexity grew 100x with AV1. CPU performance grows 5 to 10% a year. And energy regulation, especially in Europe, now puts a compliance deadline on watts per stream. Waiting does not pause any of them.
Either. Owned VPU cards in your own or existing chassis are capex. VPU instances on partner clouds are opex. Most converging companies run hybrid. The economics work in all three models because the efficiency lives in the silicon, not the deployment.
Google and Meta built their own encoding ASICs (Argos, MSVP) over decades of capital investment, and none of it is for sale. VPUs are the open market’s route to the same architecture: purpose-built silicon, available to every platform that competes openly.
One to two weeks. The pilot runs your real workloads and produces your numbers: cost per stream, watts per stream, streams per rack. The investment case is built from your data, not our claims.
Moving to VPU-native infrastructure is not a technology decision.
It is an infrastructure decision requiring internal alignment.
The Deployment Playbook is how converging companies got there.