Voice Cloning and Biometric Data Notice
Voice recordings are biometric data. What we do with a clip you upload, what consent you must have before you upload it, and how to have a voice erased.
Effective 27 July 2026
Read this before you upload a clip of anybody but yourself — and read it if you have just been told that your voice is on this platform. It is written for both of you.
01Why this document exists
A recording of someone speaking is not an ordinary file. It carries a voiceprint — a pattern specific enough to identify the person and, in this product, specific enough to reproduce them. Many legal regimes class that as biometric data and hold it to a higher standard than a name or an email address.
Voice Buddy clones voices from a single clip. That is the feature, and it is also the risk: the same twelve seconds of audio that lets you build a narrator lets somebody build a convincing fake of a person who never agreed to it. This document sets out what we do with a clip, what you must have before you upload one, and what a speaker can demand if their voice ends up here without their agreement.
02What we collect and what we derive
- The clips you upload
- Stored as files in our Google Cloud Storage bucket, under a path prefixed with your workspace, along with the file name, format, size and duration. One clip per voice is marked the reference — that is the one the engine actually conditions on.
- A preview
- A short sample of the cloned voice, generated so the voice can be auditioned in the library.
- A trained adapter
- Only for professional clones. A model file fine-tuned on your workspace’s renders of that voice. This is derived biometric data: it encodes the voice even though it is not a recording of it.
- Every render made with the voice
- The generated audio and the script that produced it, kept in the workspace’s history.
We do not run voice identification, speaker matching or emotion analysis on your audio, and we do not use it to identify anyone.
All of it is stored and processed in the United States: this application and its storage bucket in Google Cloud’s us-east4 region, the speech engine and the GPU machines that fine-tune adapters in us-central1, and the engine’s own bucket in Google’s United States multi-region. If you are outside the US, uploading a clip moves it to the US; where a transfer needs a legal mechanism we rely on the standard contractual clauses.
03The two kinds of clone behave differently
Instant cloning
Zero-shot. The engine reads your reference clip at render time and imitates it. Nothing is trained; the clip is the whole of it. Delete the voice and the clip goes with it, and the ability to reproduce that voice goes with the clip.
Professional cloning
A workspace admin starts a training run and a GPU job fine-tunes a dedicated adapter for the voice. The training set is built from that workspace’s own finished renders of that same voice — the audio paired with the script that produced it. A list of those clips, with the scripts in plain text, is written into the speech engine’s storage bucket for the training machine to read.
04Who can use a voice once it is here
- Everyone in the workspace. A cloned voice belongs to the workspace, not to the person who uploaded it. Any member can generate speech with it; viewers can hear it.
- Everyone on the platform, if it is shared.The voice library has a “Share to community” control. Turning it on makes the voice usable by every other workspace on Voice Buddy — not a chosen few, and not read-only: they can generate new speech in that voice. Do not share a voice unless the speaker has agreed to exactly that, in those words.
- Our engineers, where access is needed to operate or debug the service.
Voice clips are not sent to any third-party model provider. Cloning, synthesis, transcription and training all run on our own engine and our own infrastructure. The Privacy Policy lists the subprocessors that do receive data and what each one gets.
05How long a voice is kept, and what deletion actually removes
Nothing expires on a timer. A voice stays until somebody deletes it. Our retention schedule for voice recordings and everything derived from them is therefore event-driven rather than date-driven, and the table below is that schedule: for each artefact it names the event that destroys it, and where an artefact outlives the event you would expect to destroy it, it says so rather than rounding the answer down. We publish the schedule the system actually keeps instead of a destruction deadline nothing in the product would enforce.
If you want destruction sooner than one of those events, ask for it: we act on a request to erase a voice within 30 days of receiving it, whether it comes from the workspace that uploaded the clip or from the person who was recorded.
- Deleting a voice
- Erases every uploaded clip and the preview from object storage at once — subject only to the soft-delete tail described at the bottom of this table — then removes the voice.
- …but the renders survive
- Audio already generated with the voice stays in the workspace’s history, unlinked from the deleted voice. That audio still sounds like the speaker. Delete those renders too if that is the point of the exercise.
- …and a trained adapter is left behind
- A professional clone’s fine-tuned adapter is written by the speech engine into the engine’s own storage, outside the per-workspace area. No action in the product deletes it — not deleting the voice, not deleting the workspace. It is unreachable, because the record naming it is gone, but it exists until an operator removes it, and it encodes the voice.
- …and so is the training manifest
- The list of clips a training run learned from, described above, is a second file in the engine’s storage — outside the per-workspace area, at a name derived from the voice’s id, and containing every script in plain text. Deleting the voice does not remove it. Deleting the whole workspace does: that teardown is given the trained voices’ ids and erases their manifests as well as the workspace’s own files. Delete one voice and carry on, and its scripts stay in the engine’s bucket until an operator removes them.
- Deleting the workspace
- Removes the database rows and queues the deletion of every stored file the workspace owned, including the voice clips, plus the training manifest of every voice that was trained. The storage half runs asynchronously in pages and retries on failure, so it completes shortly after rather than instantly. The adapters are the one thing it still does not reach.
- …and deletion has a tail
- A deleted object does not vanish the instant the record of it does. Both storage buckets run a 7-daysoft-delete window, so an erased clip is recoverable by us for seven days and then is not; neither bucket keeps old versions of a file. The database rows that described it survive in Cloud SQL’s automated daily backups — 7 backups retained, plus seven days of transaction logs — and those roll off on their own. Nothing is copied to anywhere longer-lived, and a backup is restored to recover from a failure, never to bring one deleted voice back.
If you need a genuine, complete erasure of a person’s voice — the clips, the renders made with it, any adapter trained on it and the manifest of scripts that training left behind — email support@workflowcorp.com, marked for the attention of the privacy team, and ask for it explicitly. Deleting the voice in the console is not, on its own, that.
06The consent you must have before you upload
You may only clone a voice that is your own or one you have the speaker’s express, informed permission to clone. Consent should be in writing, kept, and specific. At minimum it should cover:
- Who — the speaker, identified, and who is doing the cloning.
- What — that a recording of their voice will be used to build a synthetic copy capable of saying things they never said, and that it may be fine-tuned further on audio generated from it.
- What for— the actual uses. “Marketing” is not a use; “pre-roll ads for our own products on our own channels” is.
- How long — how long the voice may be kept and used, and what happens at the end of it.
- Who else — whether the voice may be shared to the community library, where any workspace on the platform could generate with it. If you have not asked about this specifically, do not share.
- How to stop it — that the speaker can withdraw consent, how they do so, and that you will delete the voice when they do.
For a child’s voice, consent comes from a parent or guardian — the account doing the cloning belongs to an adult either way, since Voice Buddy accounts are for people aged 18 or over. For a deceased person, consent comes from whoever holds the rights to their voice and likeness. A recording you commissioned or paid for is not consent to clone the person who made it — a voice actor licenses a performance, not their voiceprint.
07If this is your voice and you did not agree
You do not need an account here, and you do not need to know which customer did it.
- Email support@workflowcorp.com with voice misuse in the subject line, which is what routes it to the people who handle abuse rather than to the billing queue. Tell us who you are, where you encountered the audio, and anything that helps us find it — a link, a file, a date, the name it was published under.
- We will acknowledge within two business days, and we can take a voice out of service — so that nothing further can be generated with it — while we investigate.
- If it is your voice and consent was not held, we will delete the voice, its clips, its preview, any adapter trained on it, and — on request — the renders made with it.
- Cloning a voice without consent is a termination-grade breach of the Acceptable Use Policy, not a warning.
You may also have statutory rights here. Several US states treat a recording of a voice as a biometric identifier in its own right — Illinois and Texas most explicitly — and a number of state privacy laws count a voiceprint as sensitive personal data carrying rights of access, correction, deletion and complaint. We do not ask which state you live in before honouring any of that. We answer a request about a voice from anyone, anywhere, within 30 days, whether or not a statute obliges us to. If you would rather take it to a regulator or a court instead, you are entitled to, and you do not have to come to us first.
08What this product does not do yet
Listed here rather than left to be discovered, because each one is a protection a reader might reasonably assume exists:
- No consent record. Nothing is captured at upload time tying a voice to a documented permission.
- No verification that a clip is of you. Any audio can be uploaded as any voice.
- No watermark on generated audio. A finished file carries no marker identifying it as synthetic or as ours, so nothing downstream can detect it.
- No automated screening. Nothing inspects a clip or a script before it renders; enforcement is reports and human review.
- No scheduled destruction. Biometric data is kept until somebody deletes it.
None of those five is a decision taken against building it. They are work not yet done, and the place to say so is here rather than in a support reply after something has gone wrong. What we do commit to today is narrower, and we hold to all of it:
- We never sell a voice. Voiceprints and the recordings behind them are not sold, rented, traded or licensed to anybody, and are not disclosed outside the service except to run it or where the law compels us.
- We do not identify people with it. No speaker recognition, no matching one recording against another, no enrolment of any voice into anything that would.
- Training never leaves the workspace.A fine-tune draws only on that workspace’s own finished renders of that one voice, and the adapter it produces is usable only by that voice. No clip trains a shared or general-purpose model, and no clip is sent to an outside model provider that could train on it.
- Consent is required before an upload, and withdrawal is acted on after it. A request to erase a voice is honoured within 30 days, whether it comes from the customer or from the person who was recorded.
- We tell you if it leaks. If voice data is exposed in a security breach we notify the affected workspaces without undue delay and within 72 hours of becoming aware of it. We hold no contact details for a speaker who is not a customer, so if you have written to us about your voice we will use the address you wrote from.