The STM32N6 is being used to run an offline speech recognition system that can turn spoken English into text without a network connection. The project runs on the STM32N6570-DK, using its onboard microphone to capture speech and its display to show the recognized words. This system is designed to recognize different English words. At the core is the STM32N6's Arm Cortex-M55 and its ST Neural-ART NPU, which is built for running neural-network models directly on the microcontroller. ST specifies the Cortex-M55 at up to 800 MHz and the Neural-ART accelerator at up to 600 GOPS.
The speech pipeline starts with the microphone audio being processed by the Cortex-M55, which converts the audio into log-mel spectrogram features. These features are then passed to CynipsNet v1.0, a speech recognition model with about 7.2 million parameters, which runs on the STM32N6's NPU. The project reports a 5.32% word error rate on the LibriSpeech test-clean set and 14.84% on test-other, using greedy decoding without a language model. The model was trained using public speech data along with 248 hours of recordings made with the STM32N6570-DK's own microphone. The project also reports about 0.2 W peak power during audio conversion and model computation, although the developer notes that power use has not yet been optimized.
The current system is focused on English speech, with output shown as capital letters and without punctuation. It works best when the speaker is around one metre from the microphone, and the developer notes that strong accents can reduce recognition quality. The project includes the source code, trained model weights and ready-to-flash firmware, making it possible to reproduce the setup on the STM32N6570-DK. This makes the project interesting not because it turns an MCU into a full phone assistant, but because it shows how an STM32 microcontroller with an integrated NPU can perform real speech recognition locally, opening the door to voice input for embedded devices where sending audio to the cloud is not practical.