BenchmarkXPRT Blog banner

Tag Archives: large language models

WebXPRT 5: AI tests now, lots of room for growth

In past blog posts, we’ve discussed our goal of developing one or more experimental WebXPRT workloads focused on local, browser-side AI technologies. While many of us regularly interact with cloud-based AI apps and services through a browser, on-device AI capabilities are growing rapidly, and we want WebXPRT to continue to evolve with them.

There are several driving factors behind that growth. Web API technologies keep maturing, giving browsers direct access to the hardware they need for real inference work. Advanced GPU and NPU technology is now widely available in consumer devices, so the local computing power necessary to run AI applications on-device is in reach for many users. And for many organizations, there are compelling reasons to execute increasingly vital work like LLM inferencing and agentic coding tasks on local machines—such as data privacy, regulatory compliance, and cost control.

The reasons for the experimental workload approach

The expansion of on-device AI is exactly the type of shift we built the experimental workload concept to capture. As we shared when we first announced the WebXPRT 5 workload lineup, an experimental workload section gives us the flexibility to put cutting-edge measurement tools in users’ hands—even if those tools won’t yet run on every platform WebXPRT has traditionally supported. Experimental scores stay separate from the main overall score and are completely optional, so we can add tests without affecting comparability or asking anyone to retest. That approach maintains WebXPRT’s strengths while preparing the benchmark for the future—and giving all of us valuable information today.

The AI functions that WebXPRT 5 measures today

WebXPRT 5 already includes four workloads that utilize AI capabilities: Video Background Blur with AI, Detect Faces with AI, Image Classification with AI, and Document Scan with AI. These workloads use machine learning—computer vision and OCR models such as a Caffe-based face detector, SqueezeNet for image labeling, and an LSTM-based OCR engine. WebXPRT’s ability to measure how well devices handle those types of workloads has real value, and it reflects the kinds of light browser-side inference tasks that have been in widespread use for a while.

We recognize, though, that there’s a clear need for more demanding local, browser-based AI workloads—especially LLM inference. We’re targeting that need with our experimental work. Like pretty much everyone else, we’re also developing in the midst of an incredibly dynamic technical environment. We want to purposefully move forward without sacrificing WebXPRT’s stability and reliability for the sake of expedience.

The main decisions we face

Choosing a Web AI framework. We’re still researching our open-source framework options, including candidates like ONNX Runtime Web, Transformers.js, MediaPipe, and TensorFlow.js. The ground here continues to shift. For example, Transformers.js v4 now supports a WebGPU backend and spans a very broad range of model architectures. So, one of our ongoing challenges is picking a durable foundation.

Choosing a web API. Of the primary options we’re investigating, WebGPU now has the broadest browser support (Chrome, Edge, and partial support in Firefox and Safari). WebNN remains the most promising option in the long term because it can directly target NPUs, but it’s still not ready for production—its W3C spec only reached Candidate Recommendation status in early 2026, and browser support outside of flagged, experimental builds isn’t there yet. Our web API outlook hasn’t changed much from before: WebGPU is the most practical path today, and WebNN may be an exciting possibility for tomorrow.

Choosing and sizing workloads. We’ll ideally find workloads demanding enough to genuinely stress new hardware, but light enough to run on slightly older gear without forcing huge model downloads or overextending the test’s runtime. The sweet spot for browser inference today tends to be small, quantized models, and memory ceilings and cold-start downloads are real constraints. Striking the right balance is another part of the challenge we’re working through.

We appreciate your patience

We’ve been talking about experimental WebXPRT AI workloads for a while. While we wish we already had everything worked out, we think the end product will be worth the wait. We appreciate your patience as we work through the details, and we’ll keep updating you here in the blog as we make progress.

As always, we’re open to suggestions. If you have ideas for a browser-based AI workload scenario, a framework or API you think we should weigh, a browser-based AI application you want us to consider, or any other related thoughts, please let us know!

Justin

Local AI and new frontiers for performance evaluation

Recently, we discussed some ways the PC market may evolve in 2024, and how new Windows on Arm PCs could present the XPRTs with many opportunities for benchmarking. In addition to a potential market shakeup from Arm-based PCs in the coming years, there’s a much broader emerging trend that could eventually revolutionize almost everything about the way we interact with our personal devices—the development of local, dedicated AI processing units for consumer-oriented tech.

AI already impacts daily life for many consumers through technologies such as such as predictive text, computer vision, adaptive workflow apps, voice recognition, smart assistants, and much more. Generative AI-based technologies are rapidly establishing a permanent, society-altering presence across a wide range of industries. Aside from some localized inference tasks that the CPU and/or GPU typically handle, the bulk of the heavy compute power that fuels those technologies has been in the cloud or in on-prem servers. Now, several major chipmakers are working to roll out their own versions of AI-optimized neural processing units (NPUs) that will enable local devices to take on a larger share of the AI load.

Examples of dedicated AI hardware in recently-released or upcoming consumer devices include Intel’s new Meteor Lake NPU, Apple’s Neural Engine for M-series SoCs, Qualcomm’s Hexagon NPU, and AMD’s XDNA 2 architecture. The potential benefits of localized, NPU-facilitated AI are straightforward. On-device AI could reduce power consumption and extend battery life by offloading those tasks from the CPUs. It could alleviate certain cloud-related privacy and security concerns. Without the delays inherent in cloud queries, localized AI could execute inference tasks that operate much closer to real time. NPU-powered devices could fine-tune applications around your habits and preferences, even while offline. You could pull and utilize relevant data from cloud-based datasets without pushing private data in return. Theoretically, your device could know a great deal about you and enhance many areas of your daily life without passing all that data to another party.

Will localized AI play out that way? Some tech companies envision a role for on-device AI that enhances the abilities of existing cloud-based subscription services without decoupling personal data. We’ll likely see a wide variety of capabilities and services on offer, with application-specific and SaaS-determined privacy options.

Regardless of the way on-device AI technology evolves in the coming years, it presents an exciting new frontier for benchmarking. All NPUs will not be created equal, and that’s something buyers will need to understand. Some vendors will optimize their hardware more for computer vision, or large language models, or AI-based graphics rendering, and so on. It won’t be enough for business and consumers to simply know that a new system has dedicated AI processing abilities. They’ll need to know if that system performs well while handling the types of AI-related tasks that they do every day.

Here at the XPRTs, we specialize in creating benchmarks that feature real-world scenarios that mirror the types of tasks that people do in their daily lives. That approach means that when people use XPRT scores to compare device performance, they’re using a metric that can help them make a buying decision that will benefit them every day. We look forward to exploring ways that we can bring XPRT benchmarking expertise to the world of on-device AI.

Do you have ideas for future localized AI workloads? Let us know!

Justin

Check out the other XPRTs:

Forgot your password?