Craft · 2 October 2026 · 5 min read
Using real photos in a monetised YouTube documentary: how licensing works
Which photos you can use in a video that earns money, what "credit required" means in practice, and why AI images are the wrong tool for real people.
You can use photos from Wikimedia Commons in a monetised video if their licence allows commercial use, and you credit them as the licence requires. Photos under CC BY, CC BY-SA or in the public domain qualify. Non-commercial (NC) and no-derivatives (ND) licences, and "fair use" claims, do not. This is general information, not legal advice.
Why not generate the pictures?
Image models cannot draw a real person: they draw a lookalike. A documentary about a real person or event built from lookalikes is misleading, and the mistakes are obvious to anyone who knows the subject. Real photos also carry information a generated image cannot: a real place, a real building, a real date.
Why not just buy the famous ones?
Press agency and sports photos are licensed per use and can cost hundreds of dollars each. A 15-minute documentary needs dozens of pictures, so that is not workable for most channels.
What "credit required" means
CC BY and CC BY-SA ask you to name the photographer and the licence. In a YouTube video the practical place is the description: a credit list grouped by photographer, with the licence. Tellavid builds that list automatically from each photo's own data and appends it to the description.
What to leave out
Maps, logos, flags and diagrams are rarely the picture you want and often have unclear licences. Photos of victims or the dead do not belong in a documentary even where a licence allows them. Paintings are fine for old subjects but look wrong in a story from the last century. Staged photos of a theatre production about the event are a last resort, not a picture of the event.
Matching the photo to the sentence
The hard part is relevance: a photo of the right person is not always the right photo for what is being said at that moment. Tellavid searches by the scene's own words (the place, the building, the event named in the sentence) before falling back to any photo of the subject, and uses each photo at most twice in a video.
See it applied to a real video.
Every example in the showcase includes the sentence that started it and every stage in between.
Open the showcase