AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB speech-recognition model designed to transcribe audio on-device using a CPU and no external dependencies. The company reports support for seven languages and publishes comparisons showing different accuracy leaders across benchmarks, alongside faster measured processing than Whisper base and Moonshine tiny v2 in its test setup.

Cactus Compute released Whistle, a speech-recognition model packaged in a 16.9 MB file that the company says runs on a CPU without dependencies and transcribes audio directly on a device. The release targets products such as phones, wearables, robots and smart-home systems, where keeping speech processing local may reduce reliance on network access and external transcription services.

Whistle accepts 16 kHz mono audio up to 30 seconds in one pass and supports English, German, French, Spanish, Italian, Dutch and Polish, according to Cactus Compute. The company says the language is detected automatically unless a user specifies it. Its browser demonstration downloads the model on first use and says that audio does not leave the device.

The model also returns word-level timestamps, with start and end times and a probability for each word. A separate embedding mode produces one encoder output row for each 80-millisecond frame without generating a transcript. The release describes optional keyword biasing, which can raise the likelihood of supplied phrases during decoding, and a silence check that can skip decoding when a clip falls below a loudness threshold.

Cactus says Whistle uses its C++ CPU engine and shares components with Needle, its existing model. Its published architecture has eight encoder blocks and a decoder whose depth can be selected when the model loads. The company reports that the first token takes 11.1 milliseconds for a 10-second clip on an Apple M4 Pro CPU; that figure is a result from its stated test setup, not a performance guarantee across devices.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that runs locally on CPUs and shares an engine with its Needle model.

Local Transcription on Small Devices

A 16.9 MB model that can process speech without a server could make transcription more practical on devices with limited storage, intermittent connectivity or strict privacy requirements. Local processing can keep recordings on-device, though users and developers still need to assess how a particular implementation handles stored audio, logs and permissions.

The release is also relevant to developers building voice interfaces. Whistle combines transcription with timestamps, embeddings and keyword biasing, while its shared engine with Needle is intended to let a single binary connect audio input to other model-driven functions. These are company-described capabilities; their usefulness will depend on accuracy, latency and resource use in the target application.

Cactus’s own results do not show one model leading every accuracy test. It reports Whistle ahead of Whisper base on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average, while Whisper base leads on TED-LIUM, AMI and the MLS average. That mixed result matters for teams choosing a model: size and speed are only part of the comparison, and performance can vary by language, accent and recording conditions.

Amazon

on-device speech to text software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Cactus Benchmarked Whistle

The release compares Whistle with Whisper base and Moonshine tiny v2 on file size, first-token time, decoding speed and word error rate. Cactus lists sizes of 16.9 MB for Whistle, 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. For ten seconds of audio on an Apple M4 Pro CPU, it reports first-token times of 11.1, 73.2 and 22.8 milliseconds, respectively, and decoding rates of 1,319, 266 and 262 tokens per second.

Those figures are vendor-published comparisons, not an independent evaluation. Cactus says each model ran through its official runtime at default settings: its own C++ engine with five beams for Whistle, OpenAI Whisper for Whisper base and the Moonshine voice runtime in non-streaming mode. The company defines first-token time as audio input to the first token, and calculates decoding speed after that point.

The benchmark notes also limit direct comparison. Whisper pads every input to 30 seconds, while Whistle’s reported first-token time changes with clip length: 5.9 milliseconds at five seconds, 11.1 at ten and 36.3 at 30 seconds. Cactus says Moonshine is English-only and that some benchmark results are unavailable because model authors did not publish them. It also cautions that the AMI figure reported for Whisper is from AMI-IHM, a different subset from the AMI results for the other two models.

“It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.”

— Cactus Compute, in its October 2, 2026 release

Amazon

voice recognition app for smartphones

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Accuracy Tests Still Needed

The release does not establish how Whistle performs across different CPUs, phones or embedded devices, or how memory use and battery consumption compare in longer-running applications. The published latency and throughput figures are tied to an Apple M4 Pro and particular runtime settings.

The available material also does not provide an independent replication of the word-error-rate results, details for every test configuration, or a complete set of scores across all listed datasets. Accuracy varies by benchmark, and some comparisons involve missing results or different dataset subsets. It remains unclear how well the model handles noisy environments, overlapping speakers, varied accents and speech outside the seven supported languages.

Cactus describes the model as open, but the release information provided here does not specify the exact license or all distribution terms. Developers should check those terms and verify the model’s behavior on their own hardware and audio before relying on it in a product.

Amazon

privacy-focused speech transcription device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Testing Across Real-World Devices

The immediate next step for developers is to try Whistle on the hardware and audio conditions they expect to support, then compare its recognition accuracy, response time and resource use with alternatives. The browser demo offers a short way to test supported languages, while production use would require checking the model package, runtime and licensing terms.

Cactus has not announced a future release date or a broader evaluation schedule in the material provided. Further independently reproduced tests, device-specific measurements and results for additional languages or longer recordings would help clarify where Whistle is most useful. Until those are available, its published speed and size figures should be read as company-reported results under specified conditions.

Amazon

small speech recognition models for wearables

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is Cactus Compute’s speech-recognition model, packaged as a 16.9 MB file and designed to run on a CPU without external dependencies, according to the company.

Which languages does Whistle support?

Cactus says Whistle transcribes English, German, French, Spanish, Italian, Dutch and Polish. It can detect the language automatically or use one specified by the user.

Does Whistle send recordings to the cloud?

The company says its browser demonstration processes audio on the device and that the audio does not leave the device. The release does not independently audit every possible integration or deployment.

Is Whistle more accurate than Whisper?

There is no single winner across the results Cactus published. The company reports Whistle ahead on several datasets, while Whisper base leads on TED-LIUM, AMI and the MLS average. The comparisons are vendor-published, and some datasets or subsets are not directly comparable.

How fast is Whistle?

On an Apple M4 Pro CPU, Cactus reports 11.1 milliseconds to the first token for ten seconds of audio and a decoding rate of 1,319 tokens per second. Those measurements use the company’s stated test setup and may differ on other hardware.

Source: hn

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Fast Charging Explained: The Science Behind Quick Battery Top-Ups

Lifting battery recharge times with advanced technology, fast charging involves intricate science that keeps you wondering how it all works.

Wi-Fi 6 Vs Wi-Fi 7: Upgrading Your Phone’s Wi-Fi Explained

Discover how upgrading from Wi-Fi 6 to Wi-Fi 7 can transform your phone’s performance and why staying current matters.

Golang proposal: container/: generic collection types

The latest Go proposal adds generic collection types to the container/ package, enhancing flexibility and type safety. Details are still emerging.

What Is Reverse Wireless Charging? (Phone Charging Another Device)

Here’s how reverse wireless charging turns your phone into a portable power source, but what are the key steps to use it effectively?