Pika Labs Introduces a Suite of Four Foundation Models for Generative Audio Covering Video Soundtracks, Music, Sound Effects, and Speech

In the rapidly expanding field of generative AI, tools that once specialized in a single medium have begun to branch into adjacent areas, reflecting a broader industry shift toward more integrated creative systems. 

Video generation platforms, for instance, have long focused on visual output, leaving sound design, music composition, and speech synthesis to separate specialized models. 

This separation has often required creators to stitch together multiple services, managing different interfaces, pricing structures, and technical constraints along the way.

Pika Labs, a company primarily known for its text-to-video and image-to-video models, released a new set of four foundation models under the name Pika Audio

These models address distinct but complementary parts of generative sound: Pika Soundtrack for turning silent or sparsely audio-equipped video into a full synchronized soundscape, Pika Music for producing complete tracks up to six minutes long, Pika SFX for generating sound effects from text descriptions, and Pika Speech for text-to-speech synthesis that supports both preset voices and cloning from short reference audio. 

The models are currently available exclusively through the Pika API Club.

Pika Soundtrack accepts a video input and produces synchronized music, speech, ambient sounds, and motion-aware effects that align with on-screen events. 

Users can provide a blank prompt and allow the model to interpret the footage, or they can supply directed instructions to emphasize, exclude, or include specific elements. 

Technical reports accompanying the release indicate that it processes audio at a rate of 0.617 seconds of wall-clock time per second of generated output and shows strong performance on semantic alignment and audiovisual synchronization benchmarks relative to comparable video-to-audio systems such as Hunyuan Foley. 

Pika Music generates songs from text prompts, lyrics, vocal references, music references, or combinations of these inputs. 

Tracks can reach six minutes in length, and the model supports iteration across genres, arrangements, and vocal styles. Generation of a 90-second piece averages around 6.21 seconds. The listed price is $0.015 per minute.

Pika SFX converts natural-language descriptions into sound effects lasting up to 20 seconds at 44.1 kHz stereo. 

It handles single events or sequences and incorporates details of material, space, perspective, timing, texture, and mood while avoiding unintended speech or music unless specified. Average end-to-end generation time is reported at 0.847 seconds, with pricing at $0.0002 per second.

Pika Speech produces expressive speech from text, allowing control over delivery characteristics such as pace, tone, and formality. 

It supports both built-in voices and cloning from a few seconds of reference audio, outputting up to five minutes of 48 kHz audio. The real-time factor is given as 0.02, meaning a minute of speech requires roughly one second of compute. Pricing stands at $0.01 per minute.

The company attributes the lower price points to internal work on training efficiency and inference optimization, including few-step generation and a streamlined inference stack. 

Relative to other publicly listed models, the claims include roughly 2 times greater cost efficiency for Soundtrack versus Hunyuan Foley, up to 20 times for SFX against various alternatives, 9 times versus ElevenLabs v3 for Speech (with additional comparisons to Cartesia, ElevenLabs Turbo, and Fish Audio), and up to 10 times for Music against comparable music generators. 

These figures are presented without additional caveats in the announcement materials.

Early public tests shared on social platforms have focused on practical applications such as adding complete audio layers to silent cinematic clips, generating dialogue and Foley in a single pass, cloning singing voices for short tracks, and producing motion-synced effects for action sequences. 

Several users noted the models' relative speed and the ability to combine them in workflows, for example using Soundtrack for overall video scoring and then refining details with SFX or Speech. 

The release is framed by Pika as part of a longer effort to lower the computational and financial barriers to high-quality generative media across modalities.

Published