export_ofoct.com
[00:00.44] This week, Google researchers published a paper describing results from an artificial intelligence (AI) tool built to create music.
[00:14.79] The tool, called MusicLM, is not the first AI music tool to launch.
[00:21.69] But the examples Google provides demonstrate musical creative ability based on a limited set of descriptive words.
[00:33.11] AI shows how complex computer systems have been trained to behave in human-like ways.
[00:42.94] Tools like ChatGPT can quickly produce, or generate, written documents that compare well with the work by humans.
[00:54.62] ChatGPT and similar systems require powerful computers to operate complex machine-learning models.
[01:05.78] The San Francisco-based company OpenAI launched ChatGPT late last year.
[01:14.82] Developers train such systems on huge amounts of data to learn methods for recreating different forms of content.
[01:26.24] For example, computer-generated content could include written material, design elements, art or music.
[01:37.39] ChatGPT has recently received a lot of attention for its ability to generate complex writings and other content from just a simple description in natural language.
[01:54.39] Google engineers explain the MusicLM system this way:
[01:59.97] First, a user comes up with a word or words that describe the kind of music they want the tool to create.
[02:09.80] For example, a user could enter this short phrase into the system: “a continuous calming violin backed by a soft guitar sound.”
[02:22.55] The descriptions entered can include different music styles, instruments or other existing sounds.
[02:31.84] Several different music examples produced by MusicLM were published online.
[02:39.82] Some of the generated music came from just one- or two-word descriptions, such as “jazz,” “rock” or “techno.”
[02:50.44] The system created other examples from more detailed descriptions containing whole sentences.
[02:58.94] In one example, Google researchers include these instructions to MusicLM: “The main soundtrack of an arcade game.
[03:10.62] It is fast-paced and upbeat, with a catchy electric guitar riff. The music is repetitive and easy to remember, but with unexpected sounds…”
[03:24.71] In the resulting recording, the music seems to keep very close to the description.
[03:32.14] The team said that the more detailed the description is, the better the system can attempt to produce it.
[03:41.17] The MusicLM model operates similarly to the machine-learning systems used by ChatGPT.
[03:50.47] Such tools can produce human-like results because they are trained on huge amounts of data.
[03:58.70] Many different materials are fed into the systems to permit them to learn complex skills to create realistic works.
[04:09.86] In addition to generating new music from written descriptions, the team said the system can also create examples based on a person’s own singing, humming, whistling or playing an instrument.
[04:27.40] The researchers said the tool “produces high-quality music...over several minutes, while being faithful to the text conditioning signal.”
[04:40.41] At this time, the Google team has not released the MusicLM models for public use.
[04:48.11] This differs from ChatGPT, which was made available online for users to experiment with in November.
[04:58.21] However, Google announced it was releasing a “high-quality dataset” of more than 5,500 music-writing pairs prepared by professional musicians called MusicCaps.
[05:15.74] The researchers took that step to assist in the development of other AI music generators.
[05:24.50] The MusicLM researchers said they believe they have designed a new tool to help anyone quickly and easily create high-quality music selections.
[05:38.32] However, the team said it also recognizes some risks linked to the machine learning process.
[05:47.35] One of the biggest issues the researchers identified was “biases present in the training data.”
[05:56.38] A bias might be including too much of one side and not enough of the other.
[06:03.28] The researchers said this raises a question “about appropriateness for music generation for cultures underrepresented in the training data.”
[06:16.03] The team said it plans to continue to study any system results that could be considered cultural appropriation.
[06:27.45] The goal would be to limit biases through more development and testing.
[06:33.30] In addition, the researchers said they plan to keep improving the system to include lyrics generation, text conditioning and better voice and music quality.
