AI that creates

arqiv✳

Using AI · 02 / AI that creates

Same trick, different kinds of “next”

Chatbots predict the next word. Most image and video generators start from static and remove noise, step by step, steered by your words. Music and voice models often predict the next tiny slice of sound. Different media, same core idea: learn patterns from huge numbers of examples, then generate something new that fits them.

What each kind of model predicts

Text models predict the next token, image models predict noise to remove, video models denoise patches of space and time, and audio models predict the next slice of sound. TextThe catsmellslikefish Image Videoframes → space-time patches Soundnext

Interactive · how image generators work

From static to picture

Most image generators use a method called diffusion. The model starts with pure random noise. At each step it predicts the noise mixed into the image and removes a little of it, nudged toward your prompt. After dozens of steps, a picture appears.

Pick a prompt

Step 0 of 50: pure noise

Honest note: a real model never sees the finished picture. It has learned from millions of captioned images what “sun,” “hills” and “sunset” tend to look like, and it predicts the noise to remove at each step. Change the starting noise and you get a different picture from the same prompt. This demo cheats by mixing toward a ready-made drawing so you can watch the process.

The generative AI family

What each kind is good at, and where it slips

Examples change fast. These are well-known tools as of fall 2026, listed to show the category, not as recommendations.

Text & code

Language models

Predict the next token, over and over.

Examples
ChatGPT, Claude, Gemini
Great for
Explaining, drafting, brainstorming, writing and fixing code
Watch for
Confident made-up facts (hallucinations), agreeing too easily

Images

Image generators

Remove noise step by step, guided by your words (diffusion), or build images piece by piece.

Examples
Google’s Nano Banana Pro, OpenAI’s GPT Image, Midjourney
Great for
Concept art, mockups, visual brainstorming
Watch for
Stereotyped defaults, odd hands and text, copying living artists’ styles

Video

Video generators

Denoise “space-time patches,” little cubes of image across several frames, so motion stays consistent.

Examples
Google Veo 3.1, Runway, Kling
Great for
Storyboards, short clips, visualizing ideas
Watch for
Physics that doesn’t quite work, realistic fakes of real people

Music & sound

Music generators

Turn sound into tokens (compressed slices of audio) and predict what comes next, or denoise audio.

Examples
Suno, Udio, Google Lyria 3
Great for
Sketching song ideas, background tracks, exploring styles
Watch for
Copyright questions, sound-alikes of real artists

Voice

Voice models

Predict speech sound from text. Some can copy a specific voice from a short sample.

Examples
ElevenLabs, voice modes in ChatGPT and Gemini
Great for
Narration, translation, accessibility
Watch for
Voice cloning without consent and phone scams

How video models keep things steady

A video is a stack of frames. The model cuts the stack into small cubes spanning space and time, called space-time patches, and denoises them together. 4 frames of video space-time patches

Video

Patches through time

If a model drew each frame separately, a dog’s spots would jump around from frame to frame. So video models cut a clip into small 3-D pieces called space-time patches: a little square of the picture across several frames. In 2024, OpenAI described its Sora model as a diffusion transformer that denoises these patches together. The transformer’s attention lets every patch “look at” every other one, which helps keep objects and motion consistent.

Newer video models also generate matching sound, such as dialogue, footsteps and music, in the same pass.

Big questions

Who owns it? Who agreed to it?

Generative AI learns from human work: photos, paintings, songs and voices. That raises real questions with no settled answers yet.

Training data

Learning from artists

Models learn from huge collections of human-made work, often gathered without asking. Courts in several countries are still deciding when that’s allowed and when creators should be paid.

Music · a real example

From lawsuits to licenses

In 2024, major record labels sued the AI music startups Suno and Udio. By late 2025, Universal and Warner had settled with Udio, and Warner with Suno, through deals for licensed AI music services. Other cases were still in court in 2026.

Consent

Your face, your voice

A few seconds of audio can be enough to clone a voice. Making fake images, videos or voices of real people without permission can hurt them, and in many places some uses, such as fake intimate images, scams and impersonation, are against the law.

Is it real?

Labels help. Checking still matters.

The industry is building labels for AI-made media. They help, but a missing label doesn’t prove something is real, because labels can be stripped off.

Content Credentials

An open standard from the C2PA, a group founded in 2021 by Adobe, Microsoft, the BBC, Intel, Arm and Truepic. It attaches a kind of “nutrition label” recording where a file came from and how it was edited.

Invisible watermarks

Google’s SynthID hides a watermark inside images, video, audio and text made by its AI. You can upload an image, a video or an audio clip to the Gemini app to check whether Google’s AI made it.

Find the source

Who posted it first? Is a trusted news outlet reporting it? A reverse image search can reveal where a picture came from.

Look closely

Check hands, text, reflections, shadows and backgrounds. Then remember that the newest tools rarely slip, so the source matters more than pixels.

In the classroom

Describe it before you generate it

The creative thinking happens before the computer turns on.

1 · Paper

Sketch and brief

Students sketch an image idea and write a picture brief: subject, setting, mood, colors and three must-have details.

2 · Generate

One screen

The teacher generates two or three images from student briefs on a single screen.

3 · Probe

What did AI assume?

Compare each result with its sketch. What did the AI add that nobody asked for? Try “a scientist” or “a nurse” and discuss the defaults.

Sources: OpenAI, “Video generation models as world simulators” (Feb 15, 2024); Google Developers Blog, “Introducing Veo 3.1 and new creative capabilities in the Gemini API” (Oct 15, 2025); Google, “Introducing Nano Banana Pro” (Nov 20, 2025); Google, “A new way to express yourself: Gemini can now create music” (Lyria 3, Feb 18, 2026); Music Business Worldwide, UMG–Udio settlement (Oct 30, 2025) and Warner Music–Udio settlement (Nov 19, 2025); TechCrunch, Warner Music–Suno deal (Nov 25, 2025); Digital Music News, Sony v. Udio (Aug 31, 2026); C2PA, c2pa.org; Microsoft News, C2PA formation (Feb 22, 2021); Google DeepMind, SynthID; Georgia Tech Polo Club, Diffusion Explainer.

Put curiosity to work

Read it. Try it. Question it.

Explore new AI research, or take a paper-first investigation into your classroom.