Kits AI Voice Cloning: How to Record a Dataset That Works
Kits AI docs give two dataset lengths, 30-60 minutes and 30 to 45. Treat 30 as the floor. Mono, one session, peaks -9db to -3db. Here is how to record it.
Your first Kits AI model came back and it doesn't sound like you. Sustained notes smear, consonants go robotic, and it drifts off pitch the moment you push it. The model is not the variable you control. The dataset is. Here is the shortest route to one that trains clean.
The path from mic to trained model
Every number, path and setting below comes from the Kits AI training docs, which spread these rules across eight pages that almost nobody reads before hitting Train. Follow the steps in order. Each one links back to the page it came from so you can check the figure yourself.
1. Set the file format before you hit record
Kits wants a lossless file: .wav or .aiff, at least 16-bit, sampled at 44.1kHz or 48kHz (per the file-quality-settings page). Set this in the project, not on export. In Ableton Live, bit depth lives under Preferences > Record > Bit Depth and sample rate under Preferences > Audio > In/Out Sample Rate. Logic Pro X keeps sample rate in File > Project Settings > Audio. Reaper puts the recording format in File > Project Settings > Media tab.
The same page warns that too small an audio buffer during recording produces "crackles, static noise, pops or dropouts." If you hear those, raise the buffer and re-record. You cannot fix them in the edit.
2. Clap-test the room
The recording-environment page gives you a five-second test: "Clap your hands sharply in the room and listen. If you hear a flutter or a prolonged echo, you've got reverb issues." The room ends up in the file alongside the voice, and the post-processing page warns that reverb and delay "can obscure the clarity of the vocals and confuse the AI model's pitch detection," which is why the same docs tell you to record somewhere dry rather than fix it later.
Kill it with soft mass: carpets, rugs, thick curtains, fabric furniture, canvas paintings, or foam tiles. Record toward the centre of the room, not in a corner or against a wall. Then silence the noise floor. The docs name the usual offenders: fluorescent lights, a refrigerator, HVAC that "varies the noise floor" as it cycles, computer fans and hard drives, plus the small stuff you forget on the mic like jangling keys or a creaking chair.
One counterintuitive note worth keeping: if you plan to belt, you want more open space, not less. Too much isolation ("if you're in a closet, booth, or if your microphone is surrounded with foam") can overload the mic capsule on loud phrases.
3. Pick a mic and lock your distance
A large-diaphragm condenser is the upgrade pick; a dynamic mic is fine on a budget or in a noisier room (from the microphones page). What Kits tells you to avoid, verbatim: laptop or phone mics, lapel mics, karaoke mics, headset mics, and USB or 1/8-inch input mics, which the docs allow "can be high quality, but most are intended for communication and not studio-quality recording." Run it through an interface with decent preamps and an XLR input.
Placement is the easiest thing on this list to get wrong without noticing. The placement page is specific: "about 2 inches from the microphone for your regular volume and 4-6 inches for louder phrases/belting. Always be closer than 12 inches." Sing straight into the capsule, not at an angle. Don't cup the mic. Put a pop filter one inch out. Skip the Kaotica-style foam ball, which the docs flag for "excessive isolation and overloaded mic capsules." Set your distance once and hold it; moving mid-phrase changes your tone from line to line, and the model learns the inconsistency.
4. Set your levels low
Two doc pages cover levels. One gives you a method, the other a number. The volume-level page: sing at the loudest level you expect to hit, watch the interface meters, and "stay around 30-50% of your meter. If you see yellow/red, turn it down. It's generally better to record too low than too loud." The datasets page states the target as a number instead: "volume peaks between -9db and -3db." Those are two different scales and the docs never equate them, so use whichever one your interface actually shows. Headroom is free. Clipping is not something you undo.
5. Record 30 to 45 minutes, in one sitting
Here the docs give two figures, and you should know both. The datasets page says "30-60 minutes of clean and varied audio." The content-guidelines page says "compile 30 to 45 total minutes of vocals recorded in a single session." Treat 30 minutes as the floor either way.
Spend it deliberately. The docs suggest "about 20 minutes of confident singing examples in the style you want to clone," then "an additional 10 minutes including examples of low notes, high notes, isolated phonemes, and sibilant sounds." The reason is blunt: "Your AI voice can only accurately reproduce what it hears in your dataset." A note or a consonant you never sang is a note the model has to guess at.
Three hard rules govern what you feed it:
- Mono only. "Voice cloning requires monophonic input, so do not include any vocal stacks or harmonies."
- One session. "Record your dataset in a single session so your AI voice can capture the exact frequency response of the target voice. Combining different recordings can reduce the accuracy of the cloning process."
- One vocal quality per pitch range. If you want grit on high notes, "don't include falsetto vocals in the same pitch range as gritty vocals." The docs don't spell out what a mixed dataset produces, only that you should settle the vocal character before you record.
For phoneme coverage, Kits publishes passages that hit every phoneme in American English, plus extra sibilant lines. One example line reads: "Chefs toss fresh fish with thick salt, chop hot peppers, and pack them swiftly into shiny dishes." The rest live on the content-guidelines page. And do not keep your mistakes: "whether it's audio glitches or singing mistakes, don't include anything you wouldn't want your AI voice to learn."
6. One edit pass, and no effects
The post-processing page lists eleven steps. The ones that move the needle:
- Use a volume rider or automation for a consistent level across the whole file, while keeping the natural dynamics inside each section.
- Then "a transparent compressor or limiter with a fast attack to smooth out the peaks within sections. Try to limit dynamic range to around 5db."
- Normalize to -3db. Never let anything touch 0 dB.
- EQ only to subtract: low-end rumble, mid muddiness, high hiss, since "subtle cuts are often enough." The goal is your natural tone, not a mix.
- No time-based effects. "Refrain from adding reverb, delay, or other time-based effects. These effects can obscure the clarity of the vocals and confuse the AI model's pitch detection."
- No hard cuts (use fades so edits don't click), no layered vocals, and no copy-pasting sections to pad the length. The model "benefits from the natural variation and imperfections of a continuous, unaltered performance."
7. Upload and train
Go to the voice page at app.kits.ai/voices and choose Clone a voice. There's an on-screen guide at app.kits.ai/voices/train that walks the steps. Press Train, and per the docs the model "will begin training and complete in accordance with time estimate on your progress bar." When it's done, hit Use Voice and listen back.
One caveat you should verify yourself rather than trust here: the docs note that "voice model creation is currently limited to our premium tier users." That line sits on the Kits Earn page, and plan terms move faster than blog posts, so check current pricing on Kits directly before you count on it. You can read the full tool profile on Kits AI for context.
If it's your own voice, Kits Earn lets you submit that clone for review, after which you earn "when users download outputs created using your voice model." Payouts run through Stripe on the 1st and 15th of each month once a voice has accrued more than $10, one voice per person, and it has to be trained through Kits cloning ("no merged or uploaded voices will be considered"). Optional, and beside the point of getting a clean model.
Where it breaks, and what the break sounds like
Kits doesn't publish a symptom-to-cause table, so read what follows as a map from each common complaint to the documented rule most likely sitting behind it, not as a diagnosis:
Smeared sustained notes, robotic overall. The dataset was too short, or you stitched it from more than one session. The single-session rule exists so the model captures one consistent frequency response. Two takes on two days give it two frequency responses to reconcile, which the docs say "can reduce the accuracy of the cloning process."
"It's not me." You left harmonies or vocal stacks in the input, or you mixed vocal qualities inside one pitch range. Both violate the monophonic rule. Kits can only clone a single voice at a time, so a stacked take clones the blend.
Wanders off pitch when pushed. Two causes. Either you never sang that range, so the model is extrapolating past what it heard, or a reverb or delay tail on the track confused its pitch detection during training. Cover the low and high notes, and keep the dataset dry.
Boxy or roomy. Reverb in the recording. The clap test would have caught it. Nothing in the Kits docs describes removing it after training, so the fix is a new dataset.
Crackle and digital grit. Levels ran too hot into clipping, or the buffer was too small while tracking. Re-record lower, at 30-50% on the meters.
Popped P's and B's. Plosives. The docs' answer is a pop filter one inch from the mic, with you at about two inches for normal volume. If instead the problem is S and SH sounds coming out mangled, check whether your dataset ever covered the sibilants the content-guidelines page asks for.
When Kits AI is the wrong tool for the job
All of this assumes you control the recording and the recording is a single clean vocal you own. Break either assumption and re-training won't save you.
If your only source is a finished song with the band still audible under the vocal, a live bounce with bleed, or a voice you don't have the right to clone, you can't meet the rules above. Kits does document a Stem Splitter and a vocal separation endpoint, so pulling a vocal out of a mix is at least possible on the platform. Nothing in the training docs endorses separated audio as dataset material, though, and the single-session, dry, mono requirements don't relax. If you cannot get one uninterrupted 30-minute session of your own dry vocal, a stock voice will serve you better than a thin clone of your own.
So skip the whole procedure if you don't own a clean, isolated, single-take source of the voice you want. Everything above is about giving Kits one honest performance to learn from. Without that, there is nothing here to fix.
Keep reading
How to Write With AI Without Sounding Like a Robot
Learn how to write with AI without sounding robotic. Master prompt engineering, the write-then-edit approach, voice preservation techniques, and tools like…
Hermes Agent Tutorial: Self-Improving AI Setup
Install and configure Hermes Agent so it remembers your workflows, builds reusable skills, and gets better at your tasks over time.
How to Clone Your Voice in ElevenLabs: A Beginner Guide
What ElevenLabs' docs actually say about instant voice cloning: the plan you need, the audio spec, the five settings, and the limits.