Riffusion logo

Riffusion

0.0
0 Reviews

Generate music from text prompts.

About Riffusion

Riffusion represents one of the most fascinating, unorthodox, and highly technical approaches to AI music generation ever developed. While almost every other audio AI trains on massive datasets of raw audio waveforms (the actual sound waves), Riffusion took a completely different approach: it utilized the open-source Stable Diffusion image generation model (the exact same AI used to generate pictures) and trained it entirely on images of audio "spectrograms." A spectrogram is a visual, 2D representation of audio frequencies over time. Riffusion proved that if an image AI can learn to draw a picture of a cat, it can also learn to draw a picture of a perfect audio spectrogram. Riffusion generates these visual spectrograms based on user text prompts (e.g., "church bells playing a jazz rhythm"), and then the platform simply plays the resulting image back as audio. This brilliant, "hacky" workaround resulted in an incredibly unique, highly experimental, and often chaotic tool that is deeply beloved by the open-source and hacker communities. Because it is based on Stable Diffusion, Riffusion inherited all of its image-editing capabilities. A user can use "Img2Img" (Image-to-Image) to take a spectrogram of a piano playing a melody, run it through the AI, and force it to sound like a distorted electric guitar. It allows for infinite looping and incredibly smooth interpolations between completely different genres (morphing a country song smoothly into a techno beat). While it may not produce the radio-ready vocal hits of Suno or Udio, Riffusion remains an incredibly powerful, deeply experimental sandbox for technical audio designers and electronic music producers.

Deployment

  • Cloud, SaaS, Web

Support

  • Email/Help Desk
  • Community Forum

Training

  • Documentation
  • Community Guides

Ideal Company Size

Small Employees

Pricing Overview

Free

LicensingSubscription
Supported LanguagesEnglish
Write a Review