One hardware decision now shapes your codec roadmap, rack economics, and stream quality for years. The VPU Buyer’s Guide gives engineering, operations, and procurement teams the data to make it correctly
NETINT VPU ECOSYSTEM · IBC 2026
AT A GLANCE
Selecting a VPU affects codec support, density, latency, power efficiency, and long-term infrastructure costs. This guide explains how to compare NETINT’s Logan and Quadra product families, evaluate throughput, quality, and power consumption, and choose the right hardware for cloud, broadcast, edge, and AI video workloads. It also includes real benchmark data, product specifications, and deployment guidance to simplify purchasing decisions.
Somewhere in your next planning cycle sits a line item that will outlive most of the decisions around it: which encoding silicon to standardize on. Choose well, and cost per stream falls while the codec roadmap opens up. Choose poorly, and you inherit a density ceiling, a gap to next-gen codecs like AV1, and a power bill that scales unsustainably.
NETINT has published the 2026 edition of its VPU Buyer’s Guide with product specifications, real-world measured performance, rate-distortion evidence, and the selection logic needed to inform a buying decision. Download the Buyer’s Guide here. This post walks through how the guide structures the decision, and previews the data a specifier should expect to see before committing a fleet.
Start with the architecture, not the SKU
NETINT VPUs are built on two proprietary ASIC generations. Codensity G4 powers the Logan family (T408 and T432) and offers fixed-function H.264 and HEVC transcoding at roughly 7 watts per chip. The Codensity G5 second-gen chip powers the Quadra VPU family and adds AV1 encoding, VP9 and JPEG decoding, on-device scaling, overlay and audio processing, AI inference engines, and peer-to-peer DMA with NVIDIA and AMD GPUs for cloud gaming and AI-driven workflows.
Figure 1. Codec support, power, throughput, and hardware features by ASIC generation.
The trap of borrowed hardware
The industry reflex, trained by a decade of cloud thinking, is to solve every scaling problem with a uniform pool of general-purpose compute. For immersive media, that reflex breaks down. If the pipeline relies on premium general-purpose GPUs to handle both complex spatial rendering and massive high-density video encoding, the total cost of ownership fractures. Expensive silicon built to understand a scene ends up spending its cycles on the routine, repetitive work of compressing one.
That is why the executive question has changed. It is no longer “Which server is fastest?” It is “Which workload belongs on which class of silicon, and how should the system be engineered so that silicon sustains its performance in production?”
This is the practical case for a third category of processor alongside the CPU and GPU: the video processing unit, or VPU.
If your roadmap includes AV1, 8K, AI-assisted encoding, or ABR ladders that need on-device scaling, specify G5 Quadra. If the workload is strictly H.264 and HEVC at matching input and output resolutions, Logan remains a dependable choice.
Note that Quadra delivers roughly four times the throughput per ASIC as Logan, with just three times the power, so watts per stream actually fall as you move to the newer architecture. Because both generations run the same FFmpeg, GStreamer, and SDK stack, they can coexist on the same server. This means a Logan fleet can be migrated to Quadra by swapping modules rather than rewriting pipelines.
Form factor is an operations decision
All four Quadra VPU models share the same G5 silicon, codec support, and video quality. What you are really choosing is how encoding capacity enters your racks.
Figure 2. The Quadra family. Same silicon and codecs; the differences are packaging, density, power, and AI capacity.
T1U slots into any U.2 NVMe bay with hot-swappable service. Ten modules per 1RU server enable 320 concurrent 1080p30 encodes per rack unit. T1A brings the same encoding and decoding density to servers and workstations without U.2 bays using PCIe and upgrades the AI computing capacity to 18 TOPS from 15 TOPS with the T1U. T2A doubles the silicon per PCIe slot and is the platform of choice for cloud gaming, where peer-to-peer DMA moves rendered frames from GPU memory directly into the encoder. T1M packs a full G5 into a 30 by 60 millimeter M.2 module drawing under 10 watts for OEM and edge designs, with AI engines disabled to keep within the M.2 power envelope.
For an operations executive, the headline number is high efficiency with ultra-high density: at full encode load a single T1U works out to roughly half a watt per 1080p30 stream, and a loaded Quadra Video Server with ten modules draws just 500 watts.
Hold latency to your glass-to-glass budget
Latency claims deserve numbers, not adjectives. The guide publishes measured encoder frame latency from NETINT’s V5.7 performance test report, captured through stock FFmpeg 7.1 in low-delay mode.
Figure 3. Average encoder latency per frame, single stream, low-delay mode. Dashed lines mark the 60 fps and 30 fps frame budgets.
Every resolution encodes in less than one frame interval: 2.8 milliseconds at 720p, under 5 milliseconds at 1080p for H.264 and HEVC, and 15 to 22 milliseconds at 4K. Two figures matter as much as the averages. Under load, 1080p latency holds below 7 milliseconds for H.264 and HEVC even with 32 concurrent sessions on one ASIC. And measured variance stays under 4 milliseconds in those loaded cases, which is what keeps worst-case frames inside an interactive budget. If you run the same flags in your proof of concept, you should reproduce these numbers; the guide documents the exact configuration.
Ask for rate-distortion curves, not adjectives
“Broadcast quality” is not a specification. The guide reproduces per-clip VMAF rate-distortion curves from NETINT’s published Akamai evaluation, run at four bitrates per clip against software and GPU anchors on identical cloud instances. Two examples show the shape of the evidence.
Figure 4. VMAF vs. bitrate, park_joy 1080p50. Quadra HEVC delivers equal quality to x265 medium at 22.5% lower bitrate.
Figure 5. VMAF vs. bitrate, old_town_cross 1080p50. Against NVENC P7, NVIDIA’s highest-quality preset, Quadra saves 41.3%.
Across the twelve-clip evaluation library, per-clip results range from x265 leading by 5.5% on one film sequence to Quadra saving 22.5% on high-detail nature content. The aggregate story favors Quadra: VMAF BD-rate advantages of 2.3% (H.264) and 6.7% (HEVC) over CPU medium presets, and 4.5% and 14.6% over NVENC P7. The guide shows six curves with full methodology so your engineers can replicate the tests, and the right final step is always a validation pass on your own content mix.
Quality that survives concurrency
A quality result at one stream tells you little about economics. The question that establishes the required server count is what happens at 10, 16, and 32 concurrent sessions per device.
Figure 6. Per-session HEVC throughput as concurrency rises, identical cloud hosts. Quadra holds roughly 2x NVENC P7 and 8x x265 medium at 32 sessions.
At 32 concurrent 1080p HEVC sessions, Quadra sustains 19.0 frames per second per session, compared with 8.5 for the GPU and 2.5 for the CPU anchor. With VPUs, host CPU cost per instance measures from negligible to 3% since the decode, scale, overlay, encode, and rate control functions all execute on Quadra. That offload is the property that lets ten VPUs share one modest host, and it is why the density economics compound.
Map the gains to your content mix
The guide’s quality section also aggregates PSNR BD-rate results across four content classes from NETINT’s V5.7.0 quality report, using same-generation software encoders at the medium preset as anchors.
Figure 6. Per-session HEVC throughput as concurrency rises, identical cloud hosts. Quadra holds roughly 2x NVENC P7 and 8x x265 medium at 32 sessions.
Low-delay workloads show tremendous gains with VPUs, as Quadra’s HEVC encoder saves 18.5% on 720p gaming content and 14.8% on 1080p60 CGI, precisely where fixed-function encoders typically fade. Dual-pass premium content shows moderate HEVC gains and breaks even with H.264. Quadra AV1 tracks x265-class efficiency in these PSNR tests while adding what x265 cannot: a royalty-free delivery codec with silicon compute and power efficiency.
An important note: encoder tuning matters; Quadra’s RDO level controls let you trade throughput headroom for compression efficiency per channel, and the guide documents the presets NETINT recommends for VOD, interactive, and maximum-density profiles. The Quadra series is highly configurable to meet specific quality, efficiency, and latency targets.
The procurement checklist behind the spec sheet
Engineering picks the module; procurement carries the risk. The guide consolidates the following diligence items:
- Validated electrical and thermal operational limits.
- Certifications by module: FCC, CE, EU, RoHS, REACH, HF, WEEE, and UL on the OEM part; EU RoHS across the Smart VPU line, with RoHS and UL approval on the servers.
- SMART health and temperature monitoring access.
- NVMe compatibility.
- Software support and integration components: FFmpeg plug-ins from n3.x through n8.0.1, GStreamer 1.22 / 1.26, a direct C API, SR-IOV, and container deployment.
- Forward and reverse compatibility between Logan and Quadra families.
The evaluation path NETINT recommends is equally procurement-friendly: start with one or two modules for integration and benchmarking, test on one of the VPU Ecosystem clouds, then scale with modules or turnkey 1RU servers (x86 or ARM) that ship preloaded with the Bitstreams control plane for code-free operation.
Get the full guide
The complete Buyer’s Guide covers what this post can only sample: full specification tables for the T1U, T1A, T2A, and T1M; consolidated electrical and environmental limits; measured latency and throughput tables; six VMAF rate-distortion curves with methodology; AI inference benchmarks across nine models; server platform specifications for the Quadra Video Servers and Quadra Mini Server; and the Bitstreams control plane, from dashboard telemetry to its REST API.
More than 200,000 NETINT VPUs are deployed worldwide and have encoded over one trillion minutes of video. The platform is proven; now the remaining work is to match it to your requirements, and that is precisely what the guide is built to do.
Download the NETINT VPU Buyer’s Guide
To arrange evaluation modules or a cloud trial, contact sales@netint.com.



