Describe an image with AI
Drop a picture. A vision model on your device writes the description — the file never leaves this tab.
Choose a picture
Starting up…
What the model wrote
Alt text
Prompt
Built from the model's own sentence: the narrative frame ("there is", "which appears to be") is stripped, what is left is split into comma-separated phrases, and the medium word in front comes from the description. Nothing like "8k, highly detailed" is added — those would be our words, not the model's.
How this page works
- One model, downloaded onceSmolVLM-256M-Instruct by Hugging Face (Apache-2.0), converted to ONNX. Your browser caches it, so the second visit is ready in a second or two — and once it is cached you can carry on in this tab with the network off.
- Your picture is scaled down, not sent anywhereThe picture is resized in the page to a single 512-pixel tile, which is what this model was trained to read — that is 64 image tokens. It is handed to a worker thread in the same tab. There is no upload step to opt out of.
- Three passes, three answersThe vision half runs once; the language half then answers three questions about the same picture — a long description, a one-sentence alt text, and a seed sentence that is rewritten into a prompt line.
| Path | First download | Description | All three |
|---|---|---|---|
| WebGPU | 194 MB | 6.0 s | 6.8 s |
| CPU (WebAssembly) | 251 MB | 11.6 s | 14.1 s |
An older laptop without WebGPU will be slower than the CPU row. Phones can run it, but they have to fetch the 251 MB CPU build first, so use Wi-Fi.
What this model cannot do
- It does not know who anyone is. No names of people, brands, landmarks or artists — a 256-million-parameter model has not memorised them, and it will invent one if pushed.
- It reads only the biggest words in a picture. Headlines on a poster usually come through; a receipt, a slide or a screenshot full of small text does not.
- It writes English. It was trained on English captions, so ask it for other languages and the sentences fall apart.
- It counts badly. "Two people" in a crowd of forty is the kind of mistake it makes, so check any number before you publish it.
Questions
Does my picture get uploaded anywhere?
No. There is no upload endpoint on this site — the page loads model weights from a CDN and then does everything in a worker thread in this tab. Open your browser's network panel, describe a picture, and you will see no request carrying your image. That is also why there is no account and no daily quota: the work runs on your hardware, so there is nothing for us to ration.
How is this different from imageprompt.org, describeimage.ai or Midjourney's /describe?
Those all send the picture to a server. As of September 2026, imageprompt.org gives five free credits a day and wants an account, describeimage.ai is login-free but rate-limited, ImageContext allows five alt texts per ten minutes and states plainly that images go to third-party AI providers, and Midjourney's /describe needs a paid subscription. This page gives you a much smaller model in exchange for no upload, no sign-in and no counter.
Why doesn't the description name the people, brands or places in my photo?
Because the model genuinely does not know them. SmolVLM-256M has about 256 million parameters — roughly a thousandth of the models behind the hosted services — and that budget goes into shapes, materials and layout, not into a memorised list of celebrities and logos. It will say "a man on a horse" rather than guessing a name, which for alt text is the right answer anyway.
Is the alt text good enough for an accessibility audit?
Treat it as a first draft you edit, not a finished string. It gets the subject and the setting right and stays under the 125-character convention, which is most of the typing. What it cannot know is why the image is on your page — a product photo in a shop needs the model number, a chart needs its finding. Those are the parts a checker will flag, and they have to come from you.
Can I run a whole folder of images through it?
Not on this page — the free version does one image at a time, and there is no queue to wait in because your own machine is the queue. Batch runs, larger images and the bigger SmolVLM-500M model are what the Pro plan is for. For a few dozen images, dropping them in one by one is honestly not slow: after the first one the model is already loaded, and each picture takes about seven seconds with WebGPU (about fourteen on the CPU path) on an M2 Max.
My machine has no WebGPU. Will this still work?
Yes. The page detects it and switches to the WebAssembly build, which runs on one CPU thread, downloads 251 MB instead of 194 MB and uses a different quantisation of the text decoder that is faster on CPUs. On an M2 Max the three answers take about 14 seconds this way, against about 7 with WebGPU; an older laptop will take longer. The status line always says which path you are on.
Can it write the description in Chinese or another language?
Not in this version. SmolVLM-256M was trained on English captions, and asking it for Chinese produces half-translated, ungrammatical sentences — we would rather leave it out than ship that. The page itself is bilingual; the model's output is English. If you need the description in another language, paste it into a translator, or use the Mac app below, which translates into seven languages.
When 256 million parameters are not enough
Image2Prompt for Mac runs JoyCaption on Apple Silicon through MLX instead. It is a roughly 9 GB one-time model download and it needs a Mac, but it writes far longer descriptions and gives you five prompt styles side by side — descriptive, Stable Diffusion, MidJourney, Booru tags and a short caption — plus translation into seven languages. Same rule as here: the pictures stay on the machine. It is a free seven-day trial and then a one-time purchase.