FEATUREDAI HOT (Curated Pool)· aihot-apiZH16:00 · 07·19
→ByteDance launches Seed Audio 1.0, a single model that generates dialogue, sound effects, and ambience end-to-end
Seed Audio 1.0 uses a universal acoustic encoder to model speech, sound effects, and ambience as one scene instead of stitching separate outputs. A single prompt controls character lines, emotion, and sound cue timing at 100ms precision. It does zero-shot voice cloning from a reference clip, generates roughly 2 minutes per pass, and can extend audio while keeping the voice consistent. It covers 20+ languages and adapts rhythm and pronunciation per language. Human evals show >90% usability in film, podcast, and short-drama scenarios, with MOS above 4 for most languages. Available now on Volcano Engine's experience center.
#ByteDance#Seed Team#Volcano Engine
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
ByteDance's Seed Audio 1.0 unifies speech, sound effects, and ambience in one model with 100ms timeline control via a single prompt.
sharp
This one's worth a look because it shifts audio generation from stitching clips to composing scenes. Normally you'd run separate models for speech, sound effects, and ambience, then manually align everything. Seed Audio 1.0 trains a universal acoustic encoder that treats all audio elements as parts of one scene, so a single prompt can specify who says what with which emotion, and when each sound cue enters—down to 100ms precision.
Zero-shot voice cloning works from a reference clip, generating roughly 2 minutes per pass with the option to extend while keeping the voice consistent. It handles 20+ languages and automatically adjusts rhythm and pronunciation per language. ByteDance's own human evals show >90% usability across film, podcast, and short-drama scenarios, with MOS above 4 for most languages.
I'd discount the numbers a bit—these are internal evals, the comparison models aren't named, and the Volcano Engine experience center just went live, so we don't have real creator feedback yet. But the direction is solid: moving audio generation control from "make a voice clip" to "orchestrate a sound scene" genuinely saves time for video dubbing, podcast production, and game localization. What's missing: API pricing and stem output. Those two will determine whether this goes into production pipelines.
HKR breakdown
hook ✓knowledge ✓resonance ✓