Getting Started
- Select a source avatar image from the library on the right, or import your own with + Import to Library.
- Type your text in the Text to Speak field, pick a voice and click Generate Speech Audio to preview it.
- Hit Generate Video — ImTalking will split your text into sentences, render each one and assemble the final video automatically.
Tips for Best Quality
- Use a square, well-lit, front-facing photo with a clean background.
- Keep individual sentences short (under 15 words).
- Enable Enhancer for sharper results — it applies GFPGAN face restoration to every frame.
- Use Still mode for a stable head position with minimal random movement.
- Pitch and Rate sliders adjust the voice before generation — use them subtly (±10–20 range).
Processing Options Explained
How many video frames are rendered at once. Higher = faster but uses more RAM. Default 2 is safe for most systems.
Upscales the output to 512×512 px after rendering. Adds clarity without re-rendering. Recommended when Enhancer is off.
Controls how expressive the lip and face movement is. 1.0 is natural. Go higher (1.2–1.5) for more dramatic expressions.
Pause Behaviour Between Sentences
ImTalking automatically inserts still-frame pauses between sentences based on the punctuation in your text:
- Period, Question mark, Exclamation mark → 1.0 second still pause
- Comma, Dash → 0.5 second still pause
What Stays on Your Computer
- All avatar images you import are stored in your app's local temp folder.
- Generated speech audio (MP3) is saved temporarily and deleted after video creation.
- All rendered videos are saved to
C:\Users\[you]\Videos\SadTalker\and never leave your device. - Project history is stored only inside the desktop app's local storage.
Network Usage
- ImTalking connects to the internet only to: download TTS voices via Microsoft Edge TTS, and check for SSL certificate updates for the local WebSocket connection.
- The webpage communicates exclusively with the desktop app running on your own machine (127.0.0.1).
- No analytics, no telemetry, no usage tracking of any kind.
Face Data
ImTalking uses SadTalker's 3D facial model (3DMM) locally. Face coefficients are extracted on-device and immediately deleted after generation. No biometric data is stored or transmitted.
Reach Us
🌐 mlapplications.com — our main website✉️ support@mlapplications.com — support & bug reportsAbout
Built by mlapplications.com · Powered by SadTalker, Edge TTS and GFPGAN · Running entirely on your local machine.
Appearance
System
About
Best Results
- Use a square, well-lit, front-facing portrait photo with a clean background for the sharpest output.
- Keep sentences short (under 15 words) — ImTalking splits your text by sentence and renders each one separately.
- Enable the Enhancer for GFPGAN face restoration — it adds processing time but significantly improves sharpness.
- Use Still mode for a stable head with minimal random movement, ideal for presentations.
Voice Tips
- Use the Pitch slider (±10–20 Hz) for subtle tone changes — large values can sound unnatural.
- Slow down dense text with a negative Rate value for clearer delivery.
- Star your favourite voices — they appear at the top of the voice list for quick access.
Batch Processing
- Click Generate Video again while a job is running to add it to the queue — it starts automatically when the current job finishes.
- The queue keeps running even if you close the UI window — you'll get a Windows notification when each video is ready.