Curated topic · 6 pieces · 1h 36m

Buyer Toolkit — Evaluate, Compare, Decide

0 of 6 read
ASSET 01 OF 06Start here
Blog · 10 min read

Measuring What Matters


How Cires21 is turning video benchmarking into a discipline worth trusting

Why Measuring Video Is Hard

Video encoding is one of those domains where complexity hides behind familiar metrics. A bitrate number seems straightforward. A VMAF score looks like it settles the quality question. But both are abstractions layered on top of a system with dozens of interacting variables, and each abstraction strips away context that matters.

Consider what happens during a live encoding session. The encoder must make frame-by-frame decisions under real-time constraints, adapting to scene changes, motion intensity, and content complexity. It must do this while managing thermal limits, power budgets, and latency requirements that shift depending on the delivery target. A benchmark that captures one of these dimensions while ignoring the others is not incomplete. It is misleading.

The Measurement Complexity Stack

Layered diagram illustrating increasing benchmarking complexity, starting from raw metrics (bitrate, frame rate, resolution, latency), moving to quality metrics (VMAF, PSNR, SSIM), efficiency metrics (energy and density), system context, and operational validity, with an arrow indicating rising complexity.

Exhibit 1: Each layer of the measurement stack introduces variables that simpler benchmarks tend to ignore.

 

Cires21’s C21 Live Encoder, which supports codecs ranging from H.264 and HEVC to AV1, VP9, and MPEG-2 across protocols including HLS, DASH, CMAF, SRT, and WebRTC, has to perform under all of these conditions simultaneously. That breadth of real-world deployment experience is precisely what led the company to question whether the industry’s measurement conventions were keeping pace with the technology they were meant to evaluate.

The Limits of Headline Metrics

The most common benchmarking pitfall is what might be called the “single-number fallacy.” An encoder achieves a VMAF score of 95 at a given bitrate. Another achieves 93. The first must be better. But this reasoning collapses under scrutiny when you consider what was left unmeasured.

Was the content representative of actual production workloads, or was it selected to flatter one architecture? Were the encoders configured with equivalent optimization effort, or did one receive weeks of tuning while the other ran with defaults? Was power consumption tracked per stream, or was the total system draw buried in a footnote? Did the quality metric capture tail performance (the worst 5% of frames that viewers actually notice), or did a generous average mask inconsistency?

KEY INSIGHT
A benchmark that answers only the question it was designed to ask may still mislead the person using it to make a different decision entirely. The gap between “what was measured” and “what the operator needs to know” is where most bad infrastructure investments originate.

This is not a theoretical concern. Cires21’s engineering team, led by CTO David Gonzalez, has documented specific cases where standard industry benchmarks produced rankings that reversed under real-world operating conditions. An encoder that won a controlled lab comparison consumed three times the power per stream in continuous 24/7 operation. A codec configuration that excelled on test clips produced unacceptable quality variance on broadcast content with frequent scene changes.

Efficiency as a System Property

One of the most important conceptual shifts in Cires21’s approach is treating efficiency not as a feature of individual components but as a property of the entire system. A chip that encodes streams at low wattage is only efficient if it maintains quality, handles the required density, and operates within the thermal envelope of a production rack. Efficiency is the relationship between all of these variables, not a number that lives on a datasheet.

This perspective becomes especially important when evaluating purpose-built silicon like VPUs (Video Processing Units) against general-purpose GPUs. The two architectures make fundamentally different engineering trade-offs, and comparing them requires a framework that accounts for those differences rather than ignoring them.

Cires21 confronted this directly when it began integrating NETINT Quadra T1U VPUs into the C21 Live Encoder. The encoder had evolved through three generations of acceleration: a CPU era where software encoding dominated, a GPU era that brought density improvements at higher power costs, and beginning in 2024, a VPU era that promised a different set of trade-offs entirely. Each transition demanded new measurement approaches because the previous generation’s benchmarks could not fairly evaluate the new architecture

Building a Better Benchmark

Rather than accepting industry convention, Cires21 developed a structured methodology for designing what it calls “fair benchmarks for heterogeneous video compute.” The framework centers on four principles that sound obvious in theory but are rarely followed in practice.

Designing Fair Benchmarks for Heterogeneous Video Compute

Diagram titled “Fair Benchmark Design (Cires21 Methodology)” showing four key components: objective (what and why to measure), content (diverse clips reflecting real workloads), controls (identical codecs and settings), and metrics (percentiles and VMAF), supported by a benchmark integrity layer covering power/thermal reality, reproducibility, and documentation.

Exhibit 2: Cires21’s four-pillar methodology for designing benchmarks that hold up under operational scrutiny.

First, objective definition. Before any test runs, the team establishes precisely what question the benchmark is designed to answer. “Which encoder is better?” is not a valid objective. “Which architecture delivers the lowest cost per stream at broadcast quality under 24/7 operation?” is. The specificity forces everyone involved to confront the trade-offs before the numbers are generated rather than after.

Second, normalization and controls. Every comparison must use equivalent codec settings, equivalent optimization effort, and equivalent operational conditions. When Cires21 benchmarked the NETINT Quadra T1U against the NVIDIA RTX 4000 Ada, both platforms received the same tuning methodology applied by engineers with comparable expertise on each architecture. The VisualON Optimizer was integrated to ensure perceptual quality optimization was consistent across platforms.

Third, content selection. Test clips must represent the diversity of actual production workloads: high-motion sports content, talking-head interviews, graphics-heavy broadcasts, and low-light footage. Using only friendly content (static scenes, low complexity) flatters every encoder and tells the operator nothing about behavior under stress.

Fourth, metric selection. Cires21 insists on percentile-based quality reporting rather than averages. A mean VMAF of 95 can hide a long tail where 5% of frames fall below 85, which is exactly the kind of artifact that viewers notice and complain about. The 5th percentile VMAF score is often a more honest indicator of viewer experience than the mean.

What the Numbers Revealed

When Cires21 applied this methodology to its VPU evaluation, the results were instructive. The joint whitepaper with NETINT, titled “More Streams, Less Power: Energy Trade-offs,” benchmarked the Quadra T1U against the NVIDIA RTX 4000 Ada across H.264, HEVC, and AV1 encoding workloads.

Cires 21 + NETINT Whitepaper: VPU vs GPU Benchmark results

Cires21 vpu vs gpu benchmark results

Exhibit 3: Benchmark results from the Cires21-NETINT whitepaper across energy, operations, and quality dimensions

The headline finding was striking: VPUs delivered 4.7 times greater energy efficiency, consuming between 0.4 and 0.7 watts per stream compared to significantly higher consumption on the GPU platform. But the methodology’s value showed in the supporting data. Quality was not sacrificed for efficiency. VMAF scores remained comparable across both architectures and all three codecs. The benchmarks included percentile distributions, not just averages, confirming that tail quality was maintained.

At the operational level, the numbers translated into concrete deployment advantages: a 75% reduction in rack space, a 50% decrease in power consumption, and a fourfold increase in encoding density. These are the metrics that infrastructure planners actually use to model total cost of ownership, and Cires21 published them alongside the raw performance data so operators could map the results to their own environments.

THE BIGGER PICTURE
Efficiency benchmarks only earn trust when they are accompanied by the full measurement context: content diversity, quality percentiles, power methodology, and documentation sufficient for an independent team to reproduce the results. Cires21’s published methodology meets this standard.

From Benchmarks to Deployment Decisions

The real test of any measurement framework is whether it survives contact with production. Cires21’s experience deploying VPU-based encoding across its customer base suggests that the benchmarks held. The C21 Live Encoder now operates in environments ranging from national broadcasting (RTVE, TV3, EITB) to sports streaming (Dorna Sports, Real Madrid) to telecommunications infrastructure (Telefonica, Cellnex Telecom), and the transition from GPU to VPU acceleration has proceeded without the quality regressions that poorly characterized technology transitions typically produce.

This matters because the video industry is entering a period of significant infrastructure renewal. The migration to AV1, the growth of live-to-VOD workflows through tools like Cires21’s C21 Live Editor, the expansion of AI-driven content processing through platforms like MediaCopilot, and the continued push toward lower latency via WebRTC (WHIP/WHEP) all place increasing demands on encoding infrastructure. Operators making five-year capital investments need benchmarks they can trust, not benchmarks that were designed to close a sale.

Common Pitfalls and How to Avoid Them

Cires21’s methodology documentation identifies several measurement traps that are worth highlighting because they recur across the industry.

The optimization parity problem: When comparing two architectures, it is common for the incumbent to receive extensive tuning while the challenger runs with minimal optimization. This produces results that reflect the effort invested in tuning, not the inherent capability of the hardware. Cires21’s protocol requires that each platform receive equivalent engineering attention, documented and auditable.

The friendly-content trap: Benchmarks that use only low-complexity test sequences overstate the performance of every encoder and obscure the differences that matter most. Production content is unpredictable, and measurement must reflect that unpredictability. Cires21 mandates a content mix that includes challenging material like sports with fast camera motion, concerts with strobe lighting, and news tickers with sharp text overlays.

The averages-vs-percentiles blindspot: A mean VMAF of 95 across a session can coexist with individual frames scoring below 80. Viewers do not experience averages. They experience the worst frames. Any quality claim that reports only mean scores without 5th or 1st percentile data is incomplete, and Cires21’s framework treats it as such.

The power-measurement gap: Many benchmarks omit power data entirely or report it at the wall level without isolating the encoding workload. Cires21’s methodology requires per-stream power measurement under sustained load, because power consumption often increases non-linearly as density rises and thermal management becomes a factor.

Measurement as a Competitive Advantage

There is an emerging argument that the organizations with the best measurement practices will make the best infrastructure decisions, and that this advantage compounds over time. Operators who invest in rigorous evaluation avoid expensive technology bets that underperform in production. They build vendor relationships grounded in transparency rather than marketing claims. They create internal knowledge that makes each subsequent decision faster and better informed.

Cires21’s trajectory illustrates this pattern. By publishing its benchmarking methodology rather than treating it as proprietary, the company has positioned itself as a trusted evaluation partner, not just a technology vendor. The video industry does not lack for performance data. What it lacks is performance data that operators can trust, contextualize, and reproduce. Cires21’s contribution is not simply another set of benchmark numbers. It is a framework for deciding which numbers deserve to be believed so operators can get the technical results they need.

This article is part of a series examining how purpose-built video processing is reshaping streaming infrastructure. Cires21 (cires21.com) is a video technology company founded in 2008 in Madrid, Spain, providing encoding, delivery, and AI-powered solutions for broadcast and streaming operators across Europe and the Americas.

 

ACCESS NOW:  ASIC-Based Transcoding
for High-volume Use Cases
Including social media, broadcast, interactive platforms, and service providers


ACCESS NOW

Open the original blog