BenchmarkXPRT Blog banner

Tag Archives: web API

WebXPRT 5: AI tests now, lots of room for growth

In past blog posts, we’ve discussed our goal of developing one or more experimental WebXPRT workloads focused on local, browser-side AI technologies. While many of us regularly interact with cloud-based AI apps and services through a browser, on-device AI capabilities are growing rapidly, and we want WebXPRT to continue to evolve with them.

There are several driving factors behind that growth. Web API technologies keep maturing, giving browsers direct access to the hardware they need for real inference work. Advanced GPU and NPU technology is now widely available in consumer devices, so the local computing power necessary to run AI applications on-device is in reach for many users. And for many organizations, there are compelling reasons to execute increasingly vital work like LLM inferencing and agentic coding tasks on local machines—such as data privacy, regulatory compliance, and cost control.

The reasons for the experimental workload approach

The expansion of on-device AI is exactly the type of shift we built the experimental workload concept to capture. As we shared when we first announced the WebXPRT 5 workload lineup, an experimental workload section gives us the flexibility to put cutting-edge measurement tools in users’ hands—even if those tools won’t yet run on every platform WebXPRT has traditionally supported. Experimental scores stay separate from the main overall score and are completely optional, so we can add tests without affecting comparability or asking anyone to retest. That approach maintains WebXPRT’s strengths while preparing the benchmark for the future—and giving all of us valuable information today.

The AI functions that WebXPRT 5 measures today

WebXPRT 5 already includes four workloads that utilize AI capabilities: Video Background Blur with AI, Detect Faces with AI, Image Classification with AI, and Document Scan with AI. These workloads use machine learning—computer vision and OCR models such as a Caffe-based face detector, SqueezeNet for image labeling, and an LSTM-based OCR engine. WebXPRT’s ability to measure how well devices handle those types of workloads has real value, and it reflects the kinds of light browser-side inference tasks that have been in widespread use for a while.

We recognize, though, that there’s a clear need for more demanding local, browser-based AI workloads—especially LLM inference. We’re targeting that need with our experimental work. Like pretty much everyone else, we’re also developing in the midst of an incredibly dynamic technical environment. We want to purposefully move forward without sacrificing WebXPRT’s stability and reliability for the sake of expedience.

The main decisions we face

Choosing a Web AI framework. We’re still researching our open-source framework options, including candidates like ONNX Runtime Web, Transformers.js, MediaPipe, and TensorFlow.js. The ground here continues to shift. For example, Transformers.js v4 now supports a WebGPU backend and spans a very broad range of model architectures. So, one of our ongoing challenges is picking a durable foundation.

Choosing a web API. Of the primary options we’re investigating, WebGPU now has the broadest browser support (Chrome, Edge, and partial support in Firefox and Safari). WebNN remains the most promising option in the long term because it can directly target NPUs, but it’s still not ready for production—its W3C spec only reached Candidate Recommendation status in early 2026, and browser support outside of flagged, experimental builds isn’t there yet. Our web API outlook hasn’t changed much from before: WebGPU is the most practical path today, and WebNN may be an exciting possibility for tomorrow.

Choosing and sizing workloads. We’ll ideally find workloads demanding enough to genuinely stress new hardware, but light enough to run on slightly older gear without forcing huge model downloads or overextending the test’s runtime. The sweet spot for browser inference today tends to be small, quantized models, and memory ceilings and cold-start downloads are real constraints. Striking the right balance is another part of the challenge we’re working through.

We appreciate your patience

We’ve been talking about experimental WebXPRT AI workloads for a while. While we wish we already had everything worked out, we think the end product will be worth the wait. We appreciate your patience as we work through the details, and we’ll keep updating you here in the blog as we make progress.

As always, we’re open to suggestions. If you have ideas for a browser-based AI workload scenario, a framework or API you think we should weigh, a browser-based AI application you want us to consider, or any other related thoughts, please let us know!

Justin

Web APIs: Possible paths for the AI-focused WebXPRT 4 auxiliary workload

In our last blog post, we discussed one of the major decision points we’re facing as we work on what we hope will be the first new AI-focused WebXPRT 4 auxiliary workload: choosing a Web AI framework. In today’s blog, we’re discussing another significant decision that we need to make for the future workload’s development path: choosing a web API.

Many of you are familiar with the concept of an application programming interface (API). Simply put, APIs implement sets of software rules, tools, and/or protocols that serve as intermediaries that make it possible for different computer programs or components to communicate with each other. APIs simplify many development tasks for programmers and provide standardized ways for applications to share data, functions, and system resources.

Web APIs fulfill the intermediary role of an API—through HTTP-based communication—for web servers (on the server side) or web browsers (on the client side). Client-side web APIs make it possible for browser-based applications to expand browser functionality. They execute the kinds of JavaScript, HTML5, and WebAssembly (Wasm) workloads—among other examples—that support the wide variety of browser extensions many of us use every day. WebXPRT uses those types of browser-based workloads to evaluate system performance. To lay a solid foundation for the first future browser-based AI workload, we need to choose a web API that will be compatible with WebXPRT and the Web AI framework and AI inference workload(s) we ultimately choose.

Currently, there are three main web API paths for running AI inference in a web browser: Web Neural Network (WebNN), Wasm, and WebGPU. These three web technologies are in various stages of development and standardization. Each has different levels of support within the major browsers. Here are basic overviews of each of the three options, as well as a few of our thoughts on the benefits and limitations that each may bring to the table for a future WebXPRT AI workload:

  • WebNN is a JavaScript API that enables developers to directly execute machine learning (ML) tasks on neural networks within web-based applications. WebNN makes it easier to integrate ML models into web apps, and it allows web apps to leverage the power of neural processing units (NPUs). WebNN has a lot going for it. It’s hardware-agnostic and works with various ML frameworks. It’s likely to be a major player in future browser-based inference applications. However, as a web standard, WebNN is still in the development stage and is only available in developer previews for Chromium-based browsers. Full default WebNN support could take a year or more.
  • Wasm is a binary instruction format that works across all modern browsers. Wasm provides a sandboxed environment that operates at near-native speeds and takes advantage of common hardware specs across platforms. Wasm’s capabilities offer web developers a great deal of flexibility for running complex client applications in the browser. Simply put, Wasm can help developers adapt their existing code for additional platforms and browser-based applications without requiring extensive code rewrites. Wasm’s flexibility and cross-platform compatibility is one of the reasons that we’ve already made use of Wasm in two existing WebXPRT 4 workloads that feature AI tasks: Organize Album using AI, and Encrypt Notes and OCR Scan. Wasm can also work together with other web APIs, such as WebGPU.
  • WebGPU enables web-based applications to directly access the graphics rendering and computational capabilities of a system’s GPU. The parallel computational abilities of GPUs make them especially well-suited to efficiently handle some of the demands of AI inference workloads, including image-based GenAI workloads or large language models. Google Chrome and Microsoft Edge currently support WebGPU, and it’s available in Safari through a tech preview.

Right now, we don’t think that WebNN will be fully out of the development phase in time to serve as our go-to web API for a new WebXPRT AI workload. Wasm and/or WebGPU appear to our best options for now. When WebNN is fully baked and available in mainstream browsers, it’s possible that we could port any existing Wasm- or WebGPU-based WebXPRT AI workloads to WebNN, which may open the possibility of cross-platform browser-based NPU performance comparisons.

All that said and as we mentioned in our previous post about Web AI frameworks, we have not made any final decisions about a web API or any aspect of the future workload. We’re still in the early stages of this project. We want your input.

If this discussion has sparked web AI ideas that you think would benefit the process, or if you have feedback you’d like to share, please feel free to contact us!

Justin

Check out the other XPRTs:

Forgot your password?