Speech Synthesis Fingerprinting: Voice List Tracking
How the SpeechSynthesis API voice list reveals your operating system and platform, and techniques to control voice-based fingerprint signals.
BotBrowser Team
Prefer the maintained product doc?
This article has a matching page in the docs center. Use the docs for the canonical setup flow, current flags, and long-term reference.
Why Voice Lists Reveal Platform Context
The short answer is that speechSynthesis.getVoices() exposes a structured view of speech services available to the current browser. Names, language tags, local-service flags, ordering, and loading behavior differ across operating systems and browser builds. A site can read those values without asking the person to play audio. That makes the list a privacy-relevant browser surface even when an application uses it only to offer a read-aloud button.
The Web Speech API was designed for useful output, not for identifying a device. The SpeechSynthesis controller accepts a SpeechSynthesisUtterance, selects an available voice when one is requested, and reports lifecycle events. The same API also permits enumeration through getVoices(). Enumeration can reveal more context than a person expects because installed language packs, system updates, browser branding, and speech-service choices all influence the returned objects.
For readers building an accessible feature, this distinction matters: a voice menu is a convenience, while a voice inventory can become a persistent observation. Keep readable text in the page and expose only the choices needed for the task. For a broader introduction to browser signals, see browser fingerprinting explained and navigator properties and platform context.
What getVoices() Exposes
Each SpeechSynthesisVoice normally includes a display name, a BCP 47 lang tag, a localService Boolean, a voiceURI, and a default flag. The specification defines these fields, but it does not require every browser to expose the same inventory or to use the same names. Two machines with the same browser version can therefore return different arrays, while two browsers on one machine can expose different subsets.
Operating-system defaults are one source of variation. Windows installations commonly use Microsoft-branded names, macOS installations expose Apple voices, and Linux configurations may expose voices supplied by speech-dispatcher or an installed provider. These examples are not fixed promises: a language pack, a system update, an enterprise image, or a user-installed speech service can add or remove entries.
Browser brand and release are another source. A browser can decide which system voices to publish, whether network-backed voices are included, and how the list is ordered. Mobile browsers and embedded web views often have their own integration choices. The useful conclusion is not that one name proves one operating system; it is that the complete set forms a context that can be compared with other declared values.
Timing adds behavioral information. On some implementations the first getVoices() call returns an empty array, followed by a voiceschanged event after initialization. Other implementations populate the array earlier. The number of events, their order, and the delay before a usable list can vary with the speech service and browser lifecycle. An application should tolerate this variability, but a measurement system can observe it.
The list can also change while a page is open. A user may install a language pack, connect a managed speech service, change the default voice, or move from an embedded view to a full browser. Code that caches an inventory forever turns a normal configuration change into stale data. Code that sends every change to analytics creates an unnecessary record of the person's device configuration.
No permission prompt is required for the usual enumeration call. That is appropriate for a voice picker, yet it means the person may not know that a page inspected the list. Treat voice metadata as contextual information. Do not use it as an identity key, do not transmit a complete array when a selected language is enough, and do not retain voiceURI values unless a documented support need exists.
Why Network Privacy Settings Do Not Change Local Voices
A VPN or proxy changes how requests reach a server. It does not replace the speech services registered with the operating system, so it cannot make two different machines return the same voice list. Network location and voice availability are separate inputs. A network privacy control can still be useful for its intended purpose, but it should not be presented as a control for this API surface.
Private browsing changes storage and session handling. It normally does not remove installed voices or alter the browser's connection to the local speech service. A site that can enumerate voices in a regular window can often enumerate them in a private window as well. The correct product decision is to limit collection and explain why a voice choice is requested, not to promise that a window mode erases the signal.
An extension or injected script can replace a JavaScript return value, but a complete replacement must also account for property descriptors, prototype behavior, event timing, and actual playback. If a page reports a voice that cannot be selected or used, the data and behavior disagree. This is why piecemeal page scripts are a poor foundation for a stable profile: they change one observation while leaving related observations untouched.
Disabling the API has a similar tradeoff. A missing speechSynthesis object, an empty list in an environment that normally exposes voices, or a permanently absent voiceschanged event can itself be unusual. Accessibility also suffers when an application needs speech but receives no usable path. A consistent workflow should distinguish a genuine unsupported environment from a configured profile and should keep text available in both cases.
The practical privacy rule is simple. Ask whether the product needs speech output, a selected voice, or only readable text. Request the smallest value that satisfies that need. Keep voice selection local when possible, state clearly when text is sent to a separate provider, and avoid using the list to infer a person, location, or device class.
How BotBrowser Keeps Profile Data Consistent
BotBrowser applies speech voice metadata from the loaded profile before page scripts run. With the profile-backed default, getVoices() reports the names, languages, local-service values, ordering, and browser-brand context recorded for the profile. A profile intended for a Windows Chrome context therefore does not inherit an unrelated Linux host inventory merely because the process is running on a Linux server.
The profile data is part of a wider identity declaration. navigator.platform, the User-Agent, locale settings, and font-related values should describe the same target context. Voice data is useful only when those declarations agree. BotBrowser's value is the coordinated profile, not a claim that every host speech subsystem behaves like the target device.
The loading pattern is included in the reported behavior. The first call may be empty when that pattern is appropriate, and the voiceschanged event follows the profile's expected sequence. This lets an application use the normal asynchronous code path instead of adding a special page script. It also makes a comparison meaningful because the data and the way it arrives are considered together.
The profile must be complete. A partial profile can omit a language, a browser-brand variant, or the ordering needed by the target context. Use the documented launch form, --bot-profile="path/to/profile.enc", and keep the profile version associated with the browser build that will run it. The --bot-speech-voices=profile setting selects profile-backed values; --bot-speech-voices=real intentionally exposes system voices and is not a protected baseline.
Third-party speech extensions remain outside the profile's control. An extension that registers another voice can change the result after the browser has started. Do not add such extensions when a stable profile comparison is required. The same caution applies to host policies that install or remove speech services during a run.
BotBrowser controls what the browser API reports. It does not promise that the host can pronounce every language, produce audio, or route sound to a particular output device. This boundary is important for automation: a profile check can validate API data and events, while a playback test must be run on the host and application combination that will actually deliver audio.
A Verification Workflow for Applications
Begin with a baseline page that records only the fields needed for the test: voice count, names, language tags, localService, default, and the order in which entries appear. Record the browser version, profile identifier, target platform, and locale beside the result. Do not upload the full array to a third party simply to make a dashboard convenient.
Use the normal asynchronous pattern. Read the list once, subscribe to voiceschanged when it is empty, and remove the listener after the first usable result or a defined timeout. Log whether the result was immediate or delayed, rather than retaining every event detail. This tests the page's compatibility without turning an optional feature into a device census.
Compare a profile run with the contract for that profile. Names should follow the declared browser and platform conventions, language tags should cover the selected locale, and local-service values should be plausible for the profile. Ordering should be stable across repeated launches with the same inputs. A different count is a reason to inspect the profile and browser version, not proof that a single field identifies a person.
Run the same check on a host operating system that differs from the profile target. The expected result is that the API still describes the profile rather than leaking the host inventory. Also run a real-system mode when the workflow requires host speech, because profile mode and real-system mode intentionally answer different questions. Keep those results in separate test records.
Exercise normal application states. Start with an empty list, dispatch a delayed voiceschanged event, remove the selected voice, change the utterance language, and report an API error. The page should keep its written content visible, explain that speech is unavailable, and let the user retry or choose another voice. These are compatibility checks, not attempts to identify an installed service.
Check cross-signal alignment. The declared platform, User-Agent, locale, fonts, and speech list should describe one coherent profile. If the voice list says one thing while another API says something incompatible, fix the profile or configuration at the source. Do not hide a mismatch by deleting the voice menu or by collecting more unrelated data.
For a Playwright run, launch the BotBrowser executable with the profile argument and evaluate the same helper in the page context. Keep the browser brand, profile file, and headless setting in the test record. Headless mode can still expose profile-backed voice data, but it does not create audio output. A successful API assertion therefore does not replace a host playback test.
Repeat the baseline after a fresh browser start and after a second page requests the list. The result should remain within the profile contract, while ordinary page state such as a selected voice may be local to each page. This separates profile behavior from application caching and catches tests that accidentally reuse a stale array from an earlier context.
Check the fields that an application actually consumes. A page that only selects by language does not need to compare every display name, whereas a compatibility report may need the full ordered set. Write the acceptance rule before collecting data, and discard fields that have no decision attached to them. A narrow rule is easier to explain to users and support teams.
When a mismatch appears, compare the launch arguments, profile date, browser brand, locale, and extension set before changing application code. A host service installed during a test can alter a real-system result without changing the profile. Keep one reproducible command for each mode so an operator can tell whether a difference came from the browser configuration or from the host.
An application can expose a useful support page without exposing a voice inventory. Show the selected language, whether a voice was ready immediately or after a change event, and the user-facing error state. Do not show or transmit unrelated voice names merely because the API made them available. This keeps diagnostics focused on the task and reduces accidental disclosure.
Support teams can still reproduce a report by recording the browser build, profile identifier, locale, mode, and visible error message. Those values explain the workflow while avoiding a permanent catalogue of local services. When a user opts into a detailed diagnostic, state what will be collected and delete it when the support case is closed.
Treat a voice-list comparison as a bounded compatibility check, not as a reason to keep a historical device record. A useful test report can say that the selected locale was available, the list arrived through the expected event path, and the fallback remained usable. It does not need every voice name or URI. This smaller report is easier to review after a browser update, safer to share with support, and less likely to become an unintended cross-session identifier.
Teams should also define ownership for profile updates. When a browser release or language-pack change alters the expected list, record who approved the new contract, which application states were exercised, and when the old result may be deleted. That change process keeps accessibility testing, privacy review, and browser configuration aligned without treating a transient API observation as permanent identity data.
Boundaries, Failure Cases, and Practical Decisions
Voice availability is not playback availability. A returned voice can have a valid name and language while the host service is unavailable, muted, or unable to pronounce a particular string. A profile can make the browser report the expected metadata, but the operating system still owns audio production. Document this distinction wherever a product promises read-aloud behavior.
Browser updates can change names, ordering, event timing, or the set of voices exposed by a brand. Keep profiles tied to supported browser versions and revalidate after a browser update or language-pack change. Do not hard-code a universal count; use the profile's recorded contract and allow a documented update when the platform changes.
Localization needs the same care. A Japanese profile should carry Japanese language coverage when that is part of the target context, but a lang tag does not translate text or guarantee pronunciation. Let the user choose an available option, keep the original text visible, and provide a clear fallback when no matching voice is ready.
Privacy and accessibility reinforce each other here. A read-aloud control should start after an explicit action, include pause and stop behavior where appropriate, and remain keyboard accessible. The text should never disappear while the page waits for voiceschanged. A person who does not want application speech should have an off switch that does not remove the written content or interfere with their screen reader.
When support needs a diagnostic, prefer a coarse result such as "profile voice list unavailable" or "selected language missing." Avoid serializing every voice name, URI, and event timestamp into routine analytics. If a customer deliberately supplies a voice list for troubleshooting, handle it as support data with a stated purpose and retention period rather than as a default telemetry stream.
The safe operating checklist is therefore: load a complete profile, pin the browser version, keep third-party speech extensions out of the run, compare the API result with the profile contract, test delayed and empty states, and test actual playback separately. BotBrowser provides the profile-backed browser values. The host and application remain responsible for audio, user controls, and accessible fallback content.
Public Sources
The W3C Web Speech API specification defines SpeechSynthesis, voices, utterances, and lifecycle events. MDN's SpeechSynthesis reference documents the browser-facing interface and the behavior applications need to handle. BotBrowser's Speech Synthesis Protection guide documents profile and voice-mode configuration.
These sources support the public boundaries in this article: voice enumeration is an API capability, implementations can vary, profile-backed data is distinct from host playback, and applications should preserve a readable path when speech is unavailable. For adjacent signals, compare font fingerprinting with the same restraint: align declared context, test the supported surface, and avoid collecting more device detail than the workflow requires.
Related Articles
Take BotBrowser from research to production
The guides cover the model first, then move into cross-platform validation, isolated contexts, and scale-ready browser deployment.