spot_img
HomeResearch & DevelopmentUnlocking Long Song Creation: A New AI Approach with...

Unlocking Long Song Creation: A New AI Approach with Editable Music Scores

TLDR: BACH is a new AI model for generating long, human-controllable songs. Unlike previous methods that struggle with raw audio, BACH uses a “compose-first, perform-later” strategy based on editable symbolic music scores, treating music in bar-level units. This approach significantly improves efficiency, song duration, and perceptual quality, outperforming existing systems like Suno in several aspects and allowing users to easily edit generated music.

Generating full-length, high-quality songs with artificial intelligence has long been considered one of the most significant challenges in music AI-generated content (AIGC). Traditional methods often struggle with four key limitations: how much control users have, how well the models can adapt to new situations, the overall sound quality, and the duration of the generated music. These issues largely stem from the common approach of trying to teach AI models complex music theory directly from raw audio, a task that proves incredibly difficult for current technologies.

A new research paper introduces a groundbreaking solution called Bar-level AI Composing Helper, or BACH. This model is the first of its kind to be specifically designed for song generation through human-editable symbolic scores. Instead of directly generating audio, BACH adopts a unique “compose-first, perform-later” strategy. This means the AI first creates a musical score, much like a human composer would, and then specialized models render this score into vocals and instruments to form a complete song. This separation simplifies the learning process for the AI and offers significant advantages for users.

The core idea behind BACH is to treat the musical bar as the fundamental building block of a song. This aligns more closely with how music theory is structured, ensuring rhythmic stability and better coordination between different musical parts. By focusing on symbolic scores, BACH allows for a more transparent and editable generation process. Users can directly modify the sheet music and lyrics, making adjustments much easier than with black-box audio generation systems.

The BACH system operates through a three-stage pipeline. First, a large language model processes a user’s request to generate multilingual lyrics along with style tags. Second, these lyrics are fed into BACH’s symbolic-score generation module, which segments the song into bars and separates the vocal and accompaniment parts, producing a human-readable score in ABC notation. Finally, the vocal and accompaniment tracks are rendered separately using tools like FluidSynth for instruments and VOCALOID for vocals, and then mixed together to create the final song.

Key innovations within BACH include its bar-level tokenization, which breaks down music into manageable 16-character patches. It also introduces Dual-NTP (Track-Decoupled Next-Token Prediction), a method that separates vocal and accompaniment tokens during generation. This prevents one track from overpowering the other and allows for independent refinement of each part. Furthermore, BACH incorporates a “Chain-of-Score” structural conditioning, similar to the “Chain of Thought” in large language models, to maintain song-level coherence across long durations, ensuring that intros, verses, choruses, and outros flow together seamlessly.

Experiments have shown that BACH sets a new standard in song generation. Despite being a smaller model, it outperforms many existing open-source and even commercial solutions, including popular systems like Suno, in various automatic evaluation metrics. Human evaluations further confirm its superiority in subjective aspects like vocal quality, melodic attractiveness, and overall musicality. Notably, BACH can generate songs that are significantly longer than those produced by other models, often spanning several minutes, and does so with remarkable efficiency, creating songs in minutes rather than hours.

Also Read:

While BACH represents a significant leap forward, the researchers acknowledge areas for future improvement. Enhancing the final audio quality, addressing the scarcity of large, high-quality open-source music datasets, and developing more accurate evaluation metrics that better reflect human perception are ongoing challenges. Nevertheless, BACH demonstrates the immense potential of combining domain-specific musical knowledge with advanced AI techniques to create truly human-controllable and high-quality long-form music. You can learn more about this research in the full paper: Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation.

Karthik Mehta
Karthik Mehtahttps://blogs.edgentiq.com
Karthik Mehta is a data journalist known for his data-rich, insightful coverage of AI news and developments. Armed with a degree in Data Science from IIT Bombay and years of newsroom experience, Karthik merges storytelling with metrics to surface deeper narratives in AI-related events. His writing cuts through hype, revealing the real-world impact of Generative AI on industries, policy, and society. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -