Session 2 of the three-part Confident with AI course with TAVIP
Confident with AI
Understanding and Creating Visual Media When You Are Blind
AI has changed what blindness means in relation to images.
For years, the usual question was: How can a blind person gain access to visual information created by somebody else? That question still matters, but it is no longer the whole story. We can now ask a second question: What can a blind person create for other people to see?
This manual addresses both.
- Section One: Understanding Visual Media is about using AI to describe images, choosing the right tool, judging how much confidence to place in an answer, and checking important details without sight.
- Section Two: Creating Visual Media is about generating comics, animations, promotional videos and branded title sequences—and about retaining creative control even when you cannot independently inspect every pixel.
This is not a list of whichever apps happen to be fashionable this month. Apps and models will change. The more durable skill is learning to match the tool and the level of checking to what you are trying to do.
The central principle is simple:
Do not ask whether AI image descriptions are trustworthy in general. Ask how much trust you need to place in this particular description, right now.
The same principle applies to creating images. You do not need perfect visual certainty before you can experiment, make something, improve it and share it. You need a workflow appropriate to the consequences of getting it wrong.
Section One: Understanding Visual Media
1. Build a Toolkit, Not an Allegiance
There is no single best visual AI app for every situation. Different tools make different trade-offs between speed, detail, privacy, model quality, accessibility and access to human verification.
Five useful kinds of tool currently belong in a blind or low-vision AI toolkit.
Be My AI
Be My AI, inside the Be My Eyes app, is particularly good at quick, conversational scene descriptions. It is useful when you want a natural account of an ordinary photograph, object or room and the cost of a small mistake is low.
Seeing AI
Seeing AI remains a practical Swiss Army knife. Its different modes can help with text, documents, products, people and general scenes. It is especially useful when you want a familiar accessibility tool rather than a general chatbot that happens to accept images.
Access AI
Access AI combines general AI assistance with the possibility of human verification. That makes it useful when a task begins as an ordinary AI query but turns out to need greater certainty.
PiccyBot
PiccyBot can use a mixture of models. Rather than placing all your confidence in one answer, you can ask several models and concentrate on the details on which they agree. This is valuable when you want stronger evidence than a single description can provide but the situation does not yet require a human.
Perspective Intelligence
Perspective Intelligence offers a Seeing-AI-style experience with an emphasis on privacy and on-device processing. This can be attractive when you want information to remain on your phone, need some features to work offline, or prefer not to send every image to a cloud service.
The important distinction is not “cloud bad, phone good.” It is a set of practical trade-offs.
- Cloud services may offer the strongest models, the richest answers and the fastest improvements.
- On-device services may offer greater privacy, offline access and more control over where your data goes.
- Mixtures of models may give you greater confidence through agreement.
- Human services remain essential where a plausible mistake could cause real harm.
Different tools suit different jobs. The real victory is having a small toolkit you can reach for without having to reconsider the entire AI industry every time you encounter an image.
A note about change: Product names, prices and features will date quickly. Treat the examples in this manual as categories as well as recommendations. What matters is knowing when you need an everyday describer, a second model, on-device privacy or a human verifier.
2. How Much Should I Trust This Description?
There is no correct level of trust for all AI-generated image descriptions. A description that is perfectly adequate for a joke on Facebook may be dangerously inadequate for a medication label.
Use a three-layer system:
- Green: good enough
- Amber: want it right
- Red: need it right
You do not climb through all three layers for every image. Most ordinary life belongs in green. The skill is recognising the handful of moments that do not.
Green: Good Enough
Use an everyday AI tool when a mistake would cost you little or nothing.
Examples include:
- understanding a photograph posted by a friend;
- getting the joke in a meme;
- browsing clothing, interiors or social-media images;
- gaining a quick impression of an unfamiliar object;
- exploring an image for interest or pleasure;
- getting the general mood of a book cover.
Be My AI, Seeing AI, Access AI, PiccyBot and Perspective Intelligence may all be suitable, depending on your needs and preferences.
The answer does not need to be perfect. If one detail is wrong, nothing important follows from it. Good enough really is good enough.
Amber: Want It Right
Use more than one model when you are learning, forming an impression or preparing to repeat information, but no serious decision depends on it yet.
A mixture-of-models approach asks several systems to examine the same image independently. Details reported by only one model should be treated cautiously. Details on which several models agree are more likely to be reliable.
Think of it as asking three people what they saw and writing down the common elements. Agreement does not prove that they are correct—models can share weaknesses—but it removes many solitary guesses and hallucinations.
Amber is appropriate when:
- you want more certainty than one model can provide;
- you are researching an image;
- you may mention the description to somebody else;
- a mistake would cause confusion or mild embarrassment;
- you want factual restraint rather than an imaginative account.
PiccyBot’s mixture option is one way to do this. You can also submit the same image to two independent tools and compare their answers yourself.
Red: Need It Right
Put a human in the loop when an error could affect:
- health;
- medication;
- food safety;
- money;
- personal safety;
- a legal or official decision;
- somebody’s identity;
- an important document;
- a factual claim carrying your name or professional reputation.
No current AI system can guarantee that it is correct. Strong models can be wrong in calm, fluent and highly convincing language. That confidence is precisely what makes an unnoticed error dangerous.
In a red situation, use Be My Eyes, Aira or a trusted person. If the question concerns specialist information, use a person qualified to assess it. A sighted stranger who can read a label is not automatically qualified to interpret medical, legal or financial evidence.
The rule is deliberately blunt:
If, before AI, you would unquestionably have asked another person because the consequences mattered, keep a person in the loop.
3. Read the Moment
A blind person cannot visually compare a description with the original image. Verification therefore cannot mean “look at it and confirm that the AI is correct.”
Instead, ask two questions.
Question One: How Likely Is This to Be Wrong?
Listen for places where the model may be reaching beyond what the image can establish.
Names versus kinds
“A dog” may be visible. “A golden retriever called Max” is not established by pixels alone. An image does not contain the dog’s name.
The same applies to statements such as:
- “These are brides.”
- “This is the company director.”
- “The photograph shows Manchester.”
- “She is a nurse.”
The claim may be correct because the model has been given surrounding text or metadata. It may also be an inference or invention. Ask where the information came from.
Text and numbers
Prices, dates, doses, platform numbers, serial codes and telephone numbers deserve special caution. Models often misread text inside images, especially when it is small, stylised, curved, partly obscured or low contrast.
Do not rely on a paraphrase when the exact characters matter. Ask the system to read the text verbatim, and verify consequential numbers through another method.
Counting and position
Models may miscount repeated objects or confuse left and right. Treat statements such as “three tablets,” “the second door on the left” and “four people are standing behind her” gently unless checked.
Charts, maps and forms
An AI may describe the subject and overall shape of a chart while getting the underlying values wrong. This is particularly risky because the data—not the atmosphere—is the purpose of the image.
Ask for exact labels, axes and values. For important data, obtain the source table or an accessible version of the document.
Feelings and judgements
Words such as “romantic,” “professional,” “ominous,” “luxurious” and “tender” describe an interpretation. They may be useful, but they are not objective facts.
Treat them as flavour rather than evidence.
The texture of confidence
“Appears to be” can be a sign of appropriate uncertainty. Flat certainty about something the system could only be guessing is more dangerous.
Fluent prose is not proof. Confidence belongs to the style of the answer, not necessarily to the quality of the evidence.
Question Two: How Much Would It Matter If It Were Wrong?
Use a simple stakes ladder:
- Low stakes: A mistake costs nothing. Accept the answer and move on.
- Medium stakes: You are learning, forming an impression or might repeat the information casually. Cross-check doubtful details.
- High stakes: The answer affects health, safety, money, identity, an official decision or a claim you intend to publish. Verify it properly.
Put the two questions together.
- Low probability of error + low stakes: Use green and continue.
- Uncertain answer + medium stakes: Move to amber and compare models.
- High probability of error + high stakes: Stop and use red-level human verification.
Almost everything is low stakes. The competency lies in noticing what is not—and in refusing to treat every meme as though it were a medication label.
4. Ways to Double-Check Without Sight
“Verify” should describe actions you can actually take.
Ask a narrower question
Open-ended descriptions tend to wander. Pointed questions restrict the task.
Try:
- “Read every piece of visible text exactly as it appears.”
- “How many people are present? Count them carefully.”
- “Describe only what is literally visible. Do not infer who the people are or why they are there.”
- “Is there a price? Read the exact price and currency.”
- “What details are unclear or partly obscured?”
Ask where the information came from
Use:
- “Are you getting that from the image itself or from text around the image?”
- “Which claims are observations and which are interpretations?”
- “Would you be able to know that without the caption?”
This is especially useful when an answer contains names, relationships, locations or intentions.
Ask another model
Give the image to an independent service. Agreement increases confidence; disagreement is itself valuable information. It tells you that at least one answer is unreliable and that the disputed detail deserves checking.
Request exact text
For any label or number, ask for transcription rather than summary. If the service has a dedicated document or OCR mode, compare its output with the conversational description.
Cross-check the surrounding context
Use evidence you can access:
- the filename;
- existing alt text;
- the caption;
- the text before and after the image;
- the page title;
- the destination of a link;
- your knowledge of the source.
Context is not cheating. Sighted people also use captions, layout, prior knowledge and the surrounding page to interpret what they see.
Spend human help where it matters
Be My Eyes, Aira, friends and colleagues are valuable resources. A risk-based approach means you do not have to ask a person about every unimportant image—and you do not hesitate when a human judgement is genuinely warranted.
5. Worked Example: Judging a Book by Its Cover
The cover of The Stranger Times by C. K. McDonnell is an excellent example because it contains factual elements, stylised text, promotional quotations and an overall mood.
A literal description
The cover has a red background patterned with small black dots. In the centre is a large black-and-white illustration of an old-fashioned clear glass milk bottle. Inside the lower part of the bottle is a small sailing ship on dark, curling waves. Near the top of the bottle, a ribbon-like strip curls around the neck.
The title, THE STRANGER TIMES, dominates the centre of the design in large, distressed white capital letters with heavy black outlines. The title overlaps the bottle and extends beyond its sides. The author’s name, C. K. McDONNELL, appears in white capitals along the bottom.
At the top is the line:
“What if the weird news is the real news?”
Four short review quotations surround the upper part of the bottle:
- “EXTREMELY FUNNY” — Daily Mail
- “RIPPING ENTERTAINMENT” — The Times
- “A PRATCHETTESQUE ROMP” — Guardian
- “GENUINELY FUNNY” — SFX
A small white oval emblem containing a black penguin appears near the upper-right corner.
What can we reasonably infer?
The distressed lettering, bottle, ship, red-and-black palette and exaggerated review quotations suggest comic fantasy, strangeness and perhaps supernatural or tabloid-style adventure. The cover feels energetic, eccentric and knowingly old-fashioned.
Those are interpretations, not literal facts. A different reader might call the same design chaotic, gothic, playful or retro.
What can the cover establish?
With reasonable checking, the cover can establish:
- the title;
- the author;
- the publisher’s emblem;
- the visible illustration;
- the printed review quotations;
- the broad visual style.
It cannot establish:
- whether the novel is good;
- whether you personally will find it funny;
- whether the reviewers’ comparisons are fair;
- the precise plot;
- whether the book is accessible in your preferred format.
Choosing the trust layer
If you are browsing for pleasure and want to know what atmosphere the publisher is selling, this is green. Enjoy the description and continue.
If you are writing a review and want to discuss the cover accurately, use amber. Ask another model to transcribe the text and independently describe the central illustration.
If you are reproducing a quotation in published work, checking the ISBN or making a commercial decision based on the edition, move to red for the relevant details. Verify them against an authoritative textual source or ask a person to confirm the exact cover.
Questions worth asking
- “Read all the text on the cover verbatim.”
- “Separate the title, author, tagline and review quotations.”
- “Describe the central illustration without interpreting what it means.”
- “Which parts of your answer describe visible facts, and which describe the mood?”
- “Is any text too small or unclear to read confidently?”
The point is not to defeat the cover like an examination question. It is to get the kind of information you need for the decision you are actually making.
Section One: Quick Reference
| If a mistake would… | Layer | What to use |
|---|---|---|
| Cost nothing: a meme, social photograph, mood or general impression | Green — good enough | One everyday AI description tool |
| Mislead you while learning or exploring, but cause no serious harm | Amber — want it right | A second model or mixture of models |
| Affect money, health, safety, identity, an official decision or your name on a claim | Red — need it right | A human verifier or authoritative source |
Before relying on a description, ask:
- How likely is this detail to be wrong?
- What would happen if it were wrong?
- Is the AI reporting the image, using surrounding context or making an inference?
- Do I need the general meaning or an exact fact?
- Would another model or a human be proportionate here?
Section Two: Creating Visual Media
6. From Accessing Images to Making Them
Blind people have traditionally been positioned at the end of the visual-media pipeline. Somebody else designs the cover, draws the comic, shoots the video or creates the animation. Accessibility is then added so that we can receive some account of what was made.
Generative AI changes our position in that pipeline.
A blind creator can now:
- develop an idea in words;
- turn that idea into a detailed prompt;
- generate a still image;
- ask an AI to describe the result;
- identify obvious errors through targeted questions;
- revise the prompt and generate another version;
- turn a still image into animation or video;
- add dialogue, music and sound effects;
- create promotional and branded material;
- use human review where the stakes justify it.
This does not give a blind creator perfect knowledge of the visual result. It gives her something historically scarce: control of the process.
Visual judgement is no longer always a prerequisite for participation. It can become an optional layer of refinement.
7. A Practical Creation Workflow
Step 1: Decide what the visual must do
Begin with purpose, not appearance.
Ask:
- Who is it for?
- Where will it appear?
- What must the audience understand?
- What mood should it create?
- Is any wording, logo or factual detail essential?
- What would count as failure?
A social experiment, a personal comic and a paid advertisement require different standards of checking.
Step 2: Describe the concept
Write a plain-language brief containing:
- subject;
- setting;
- action;
- composition;
- mood;
- visual style;
- colour preferences;
- required text;
- details that must not appear.
You can ask a language model to turn this brief into a generation prompt, but retain the original brief. It is your statement of intention and the standard against which later descriptions can be compared.
Step 3: Generate the first version
Use an image-generation model to create one or more candidates. At this stage, variety is often more valuable than polishing a single attempt.
Step 4: Return the result to a describing model
Do not ask only, “Is this good?”
Ask:
- “Describe the complete image literally.”
- “Read every word exactly.”
- “List anything that appears malformed, duplicated or inconsistent.”
- “Does the image contain all the elements in my brief?”
- “What might a sighted viewer misunderstand?”
- “Describe the composition and visual hierarchy.”
The model’s answer is evidence, not a guarantee.
Step 5: Compare intention with output
Check the description against your brief.
- Are all required elements present?
- Has the system introduced an unwanted character, object or implication?
- Is the wording correct?
- Does the reported mood match the intention?
- Are important objects prominent or buried?
Step 6: Revise deliberately
Do not merely ask for “better.” Identify the change:
- remove an object;
- correct wording;
- simplify the background;
- make the main character more prominent;
- preserve a logo;
- change the emotional tone;
- improve continuity between panels;
- reduce visual clutter.
Step 7: Choose the right verification layer
Use the same green, amber and red system as you would for descriptions.
- Green: personal experiments, playful images and disposable social posts.
- Amber: public creative work carrying your name, where errors would be embarrassing or undermine the work.
- Red: commercial campaigns, sensitive representations, safety information, client work or anything where a visual mistake could cause material harm.
Step 8: Publish accessibly
Creating visual media does not end with sighted viewers.
Where possible, provide:
- concise alt text;
- a fuller description for complex images;
- captions or transcripts for video;
- accurate dialogue and sound information;
- accessible versions of any text embedded in the visual.
A blind creator is particularly well placed to understand that access should be part of the work rather than an apology added afterwards.
8. Creating a Comic When You Are Blind
My third AI-generated comic concerned two theatrical, crime-fighting cats: Mr Monsieur and Mademoiselle Miss.
The workflow was straightforward:
- I developed the story and generated the visual prompt with ChatGPT.
- I gave the prompt to Google’s Nano Banana Pro to generate the comic.
- I returned the resulting image to ChatGPT for a complete description.
- I used Gemini Pro as an additional quality check.
The finished comic ran for six pages.
The story as described back to me
On page one, Mr Monsieur makes a grand entrance across a living-room “stage,” tail high and chest forward. A close-up presents his gloriously unimpressed face. Mademoiselle Miss stretches across an armchair like a dancer preparing for curtain-up. Ordinary domestic life is being interpreted through the cats’ theatrical imagination.
On page two, a household lamp becomes a spotlight. It glows, dims and dies. Mr Monsieur reacts with shocked indignation while Mademoiselle Miss appears pleased with herself.
On page three, he investigates with enormous importance: checking the curtain, patrolling the furniture and interrogating a dust bunny as though it might confess. Miss half helps and half teases.
Page four provides a red herring. A laser dot appears on the carpet, prompting a heroic but misplaced pounce.
On page five, Mademoiselle Miss discovers the cable. She taps and fiddles with it until the lamp returns to life. She gives Mr Monsieur an affectionate nudge while he attempts to preserve his dignity.
Page six is the curtain call: warm golden light, the lamp restored, Mr Monsieur posing like a triumphant leading man and Mademoiselle Miss reclining beside him with complete confidence.
What the checking found
Gemini’s quality check identified two spelling mistakes across six pages. That does not prove everything else was correct, but it turned a vague question—“Does this look all right?”—into specific, actionable information.
For a playful personal comic, this is an amber workflow:
- generate;
- obtain a structured description;
- compare it with the script;
- ask a second model to look for continuity, spelling and obvious defects;
- decide whether any remaining uncertainty matters.
The significance is larger than one cat story. Blind people can increasingly read descriptions of graphic novels, but we can also create visual stories and share them with sighted friends.
9. Turning the Comic into an Animation
The next step was to give the final scene of the six-page comic to Google’s Veo 3.1 model.
The instruction was to:
- animate the scene;
- add sound effects where necessary;
- add music;
- give the cats voices matching the look and feel of the comic.
The resulting sequence
According to the generated description, the animation became a theatrical curtain call on a wooden stage framed by heavy red velvet curtains. A floor lamp stood in the background while a bright spotlight illuminated the centre.
The large, fluffy black cat stood on his hind legs wearing a grey Elizabethan-style ruff. The smaller black-and-white tuxedo cat began lying down and then sat up to join him.
The larger cat waved a paw authoritatively. As the music swelled, he raised it in a dramatic farewell. The smaller cat copied the gesture.
The sequence opened with applause. A deep theatrical voice announced:
“The case is closed, ladies and gentlemen.”
He then called:
“Goodnight!”
An upbeat orchestral finale ended with a brass crescendo.
Why sound changes the workflow
An animated visual is not necessarily wholly inaccessible to its blind creator. Dialogue, music, timing and sound effects provide information that can be experienced directly.
This creates a hybrid form of control:
- the visual model generates the motion;
- another model describes the visual result;
- the creator directly hears the voices, music and sound design;
- the creator revises the piece through language.
I cannot personally certify every visual detail of the animation. For this experiment, that does not matter. The purpose is not to make a safety-critical film or a client advertisement. It is to create, learn and discover what has become possible.
The work may not be perfect yet. It is already possible—and now it is a matter of the tools becoming better enough.
10. Creating a Promotional Video
The promotional-video project began with a logo for The Blind AI blog.
Later, I used a newer model to generate a synthetic photo-shoot image: a male cyclist competing in a road race while wearing The Blind AI branded clothing.
About a year after that, I gave the same still image to a video-generation model and asked it to turn the photograph into a short clip from the race.
The video as described
The cyclist rides uphill on a paved road during an outdoor event in a mountainous landscape. He is viewed from behind wearing a yellow helmet and a navy-blue jersey.
The back of the jersey carries a large, stylised eye graphic. Beneath it, THE BLIND AI BLOG appears in bold white capitals. The rider also wears dark cycling shorts and white socks.
Steep rocky peaks rise in the background. Spectators line the road, cheering and waving signs. Other cyclists are visible farther ahead. Bright sunlight produces sharp shadows while the camera follows steadily behind the rider, capturing both the effort of the climb and the energy of the crowd.
The point is the pipeline
I cannot independently decide whether the colours “pop,” whether the camera movement feels cinematic or whether a marketing professional would alter the framing.
Those judgements once determined whether I could participate at all. Now they are possible refinements to a process I can initiate and control.
For the first time, a blind person can create an entire visual promotional pipeline:
- branding;
- a synthetic photo shoot;
- branded clothing within the image;
- motion generated from the still;
- a description of the finished clip;
- revisions expressed in language.
No camera, production team or sighted intermediary is required to make the first complete version.
This is not AI “doing creativity for us.” It is the quiet collapse of an old asymmetry. Visual media—once scarce, expensive and exclusionary—is becoming cheap, fast and endlessly revisable.
When that happens, access stops being only about perfect output. It becomes about control.
11. Creating a Title Ident
A title ident is a short piece of animation and sound that gives a programme, channel or project a recognisable identity.
For AI with the Blind, I turned a static cover image into an eight-second animated sequence.
The ident contained:
- a golden wireframe sphere rotating in a starfield;
- a flare of light;
- a clean voice saying, “AI… with the blind.”
The animation itself was small. What it represented was much larger.
Blindness has often been presented through soft charity colours, apologetic language and gentle accessibility icons—visual choices designed to make disability appear harmless and easy to ignore.
I wanted the opposite:
- the visual and sonic language of modern technology;
- sophistication rather than apology;
- a broadcast identity rather than a service notice;
- a signal that blind people are participants in the future of AI.
The describing AI read the sequence as inspirational, sophisticated and technological—almost like the opening of a science-fiction programme about new forms of intelligence. That was the intended effect.
A title ident turns a collection of posts into something resembling a channel: a place with its own voice, mood and expectations. It gives the project a pulse.
The ident therefore closes the creative progression in this manual:
- generate a still image;
- tell a multi-page visual story;
- animate a scene;
- create promotional motion;
- establish a complete visual and sonic identity.
We are not merely asking to be included in somebody else’s visual world. We can increasingly build the thing ourselves.
12. Creative Quality Checks
Before publishing AI-generated visual work, ask questions appropriate to the medium.
Still images and covers
- Is the main subject present and prominent?
- Is all embedded text correct?
- Are hands, faces, limbs and repeated objects plausible?
- Does the described mood match the brief?
- Has the system introduced stereotypes or unwanted symbolism?
Comics
- Are characters consistent across panels and pages?
- Is the reading order clear?
- Does each panel represent the intended action?
- Are speech bubbles attributed correctly?
- Is spelling accurate?
- Do visual details contradict the script?
Animation and video
- Does movement follow naturally from the starting image?
- Do objects or characters change unexpectedly?
- Is the spoken wording correct?
- Do the voice, music and sound effects match the intended tone?
- Does the clip contain important visual information that needs audio description?
- Are captions or a transcript required?
Branding and promotional work
- Is the name or logo reproduced correctly?
- Are colours and symbols consistent with the brand?
- Could the imagery create an unintended association?
- Is the work polished enough for its audience and purpose?
- Does it need review by a designer or intended audience member before publication?
Decide what “finished” means
Perfection is not the only legitimate standard.
A personal experiment may be finished when it delights you. A public artwork may be finished after two models agree that its key features are intact. A commercial campaign may require a professional human review.
The creator should decide the standard deliberately rather than inheriting the most cautious possible rule for every situation.
Conclusion: Independence with Judgement
In roughly a decade, blind people have moved from scraps of machine-generated recognition to conversational descriptions, mixtures of frontier models, human verification on demand and tools that generate complete visual scenes and videos from language.
That change is astonishing.
The danger is not simply that AI can be wrong. The deeper danger is being told that there is only one responsible way to use it: one approved app, one universal standard of certainty, or an all-or-nothing choice between blind trust and total refusal.
Different moments require different levels of confidence.
- Use fast, independent AI for the ordinary images that make up everyday life.
- Compare models when you want greater confidence.
- Use people and authoritative sources when consequences are serious.
- Create freely when experimentation is the point.
- Add stronger quality control as the audience, stakes and permanence increase.
Blind people should not have to choose between independence and judgement. A layered approach gives us both.
We can understand more of the visual world without pretending that AI is infallible. We can create visual work without pretending that we know every pixel. We can ask for help where it genuinely matters without making dependence the price of participation.
The visual layer is no longer an absolute gatekeeper.
That does not mean the work is finished. It means the work has finally begun.