Using AI · 02 / AI that creates
Same trick, different kinds of “next”
Chatbots predict the next word. Most image and video generators start from static and remove noise, step by step, steered by your words. Music and voice models often predict the next tiny slice of sound. Different media, same core idea: learn patterns from huge numbers of examples, then generate something new that fits them.
What each kind of model predicts
Interactive · how image generators work
From static to picture
Most image generators use a method called diffusion. The model starts with pure random noise. At each step it predicts the noise mixed into the image and removes a little of it, nudged toward your prompt. After dozens of steps, a picture appears.
Pick a prompt
Honest note: a real model never sees the finished picture. It has learned from millions of captioned images what “sun,” “hills” and “sunset” tend to look like, and it predicts the noise to remove at each step. Change the starting noise and you get a different picture from the same prompt. This demo cheats by mixing toward a ready-made drawing so you can watch the process.
The generative AI family
What each kind is good at, and where it slips
Examples change fast. These are well-known tools as of fall 2026, listed to show the category, not as recommendations.
Text & code
Language models
Predict the next token, over and over.
- Examples
- ChatGPT, Claude, Gemini
- Great for
- Explaining, drafting, brainstorming, writing and fixing code
- Watch for
- Confident made-up facts (hallucinations), agreeing too easily
Images
Image generators
Remove noise step by step, guided by your words (diffusion), or build images piece by piece.
- Examples
- Google’s Nano Banana Pro, OpenAI’s GPT Image, Midjourney
- Great for
- Concept art, mockups, visual brainstorming
- Watch for
- Stereotyped defaults, odd hands and text, copying living artists’ styles
Video
Video generators
Denoise “space-time patches,” little cubes of image across several frames, so motion stays consistent.
- Examples
- Google Veo 3.1, Runway, Kling
- Great for
- Storyboards, short clips, visualizing ideas
- Watch for
- Physics that doesn’t quite work, realistic fakes of real people
Music & sound
Music generators
Turn sound into tokens (compressed slices of audio) and predict what comes next, or denoise audio.
- Examples
- Suno, Udio, Google Lyria 3
- Great for
- Sketching song ideas, background tracks, exploring styles
- Watch for
- Copyright questions, sound-alikes of real artists
Voice
Voice models
Predict speech sound from text. Some can copy a specific voice from a short sample.
- Examples
- ElevenLabs, voice modes in ChatGPT and Gemini
- Great for
- Narration, translation, accessibility
- Watch for
- Voice cloning without consent and phone scams
How video models keep things steady
Video
Patches through time
If a model drew each frame separately, a dog’s spots would jump around from frame to frame. So video models cut a clip into small 3-D pieces called space-time patches: a little square of the picture across several frames. In 2024, OpenAI described its Sora model as a diffusion transformer that denoises these patches together. The transformer’s attention lets every patch “look at” every other one, which helps keep objects and motion consistent.
Newer video models also generate matching sound, such as dialogue, footsteps and music, in the same pass.
Big questions
Who owns it? Who agreed to it?
Generative AI learns from human work: photos, paintings, songs and voices. That raises real questions with no settled answers yet.
Training data
Learning from artists
Models learn from huge collections of human-made work, often gathered without asking. Courts in several countries are still deciding when that’s allowed and when creators should be paid.
Music · a real example
From lawsuits to licenses
In 2024, major record labels sued the AI music startups Suno and Udio. By late 2025, Universal and Warner had settled with Udio, and Warner with Suno, through deals for licensed AI music services. Other cases were still in court in 2026.
Consent
Your face, your voice
A few seconds of audio can be enough to clone a voice. Making fake images, videos or voices of real people without permission can hurt them, and in many places some uses, such as fake intimate images, scams and impersonation, are against the law.
Is it real?
Labels help. Checking still matters.
The industry is building labels for AI-made media. They help, but a missing label doesn’t prove something is real, because labels can be stripped off.
Content Credentials
An open standard from the C2PA, a group founded in 2021 by Adobe, Microsoft, the BBC, Intel, Arm and Truepic. It attaches a kind of “nutrition label” recording where a file came from and how it was edited.
Invisible watermarks
Google’s SynthID hides a watermark inside images, video, audio and text made by its AI. You can upload an image, a video or an audio clip to the Gemini app to check whether Google’s AI made it.
Find the source
Who posted it first? Is a trusted news outlet reporting it? A reverse image search can reveal where a picture came from.
Look closely
Check hands, text, reflections, shadows and backgrounds. Then remember that the newest tools rarely slip, so the source matters more than pixels.
In the classroom
Describe it before you generate it
The creative thinking happens before the computer turns on.
1 · Paper
Sketch and brief
Students sketch an image idea and write a picture brief: subject, setting, mood, colors and three must-have details.
2 · Generate
One screen
The teacher generates two or three images from student briefs on a single screen.
3 · Probe
What did AI assume?
Compare each result with its sketch. What did the AI add that nobody asked for? Try “a scientist” or “a nurse” and discuss the defaults.
Sources: OpenAI, “Video generation models as world simulators” (Feb 15, 2024); Google Developers Blog, “Introducing Veo 3.1 and new creative capabilities in the Gemini API” (Oct 15, 2025); Google, “Introducing Nano Banana Pro” (Nov 20, 2025); Google, “A new way to express yourself: Gemini can now create music” (Lyria 3, Feb 18, 2026); Music Business Worldwide, UMG–Udio settlement (Oct 30, 2025) and Warner Music–Udio settlement (Nov 19, 2025); TechCrunch, Warner Music–Suno deal (Nov 25, 2025); Digital Music News, Sony v. Udio (Aug 31, 2026); C2PA, c2pa.org; Microsoft News, C2PA formation (Feb 22, 2021); Google DeepMind, SynthID; Georgia Tech Polo Club, Diffusion Explainer.
Put curiosity to work
Read it. Try it. Question it.
Explore new AI research, or take a paper-first investigation into your classroom.