Runs on your own machine

Depth map from a photo

Drop in a photo and get a depth map back in seconds. Depth Anything V2 runs inside this tab — the picture never leaves your machine.

Depth Anything V2 Small · 94 MB on WebGPU, 26 MB on CPU · Apache-2.0 · downloaded once, then cached

Make a depth map

Drop a photo here

JPEG, PNG or WebP — or paste one with Ctrl / ⌘ + V

No photo handy? Run one of these:

One photo at a time, up to 12 megapixels and 32 MB. Anything longer than 2048 px on its long edge is scaled down first.

How it goes

  1. 1

    Bring a photo

    Drag it onto the page, paste it, or pick a file. The photo is decoded in the tab; no request carries it anywhere.

  2. 2

    Wait a few seconds

    The first visit spends 26–94 MB fetching the model. Every run after that is pure compute: 0.3 s on WebGPU, about 5.5 s on CPU for a 1280 px photo.

  3. 3

    Take the PNG

    8-bit for ControlNet and Photoshop, 16-bit when you are going to push the levels around, GIF if you just want to show the parallax.

Questions people actually ask

What is a depth map actually good for?

Three things, mostly. It is the conditioning image for ControlNet's depth models in Stable Diffusion, so your generation keeps the geometry of the original shot. It drives 3D-photo and parallax effects — the wobble preview on this page is the same idea. And it is a selection mask for depth-dependent grading: fake bokeh, atmospheric haze, relighting the background only.

How accurate is it, and what are the units?

There are no units. Depth Anything V2 predicts relative inverse depth: larger values are nearer, and the numbers for this page's street sample land between 0.00 and about 7.3. Two different photos are not on a common scale, and nothing here measures metres. Within one photo the ordering is reliable — reflective surfaces, glass and blown-out sky are where it is least sure.

Does my photo get uploaded?

No. The only thing this page downloads is the model file, and it never posts anything back. Open your browser's network panel and run a photo through: after the weights are cached you will see zero requests. That is also why the tool works with the network switched off.

What is the 16-bit PNG for?

8-bit gives you 256 depth steps, which is enough to look at and enough for ControlNet, but it bands badly the moment you stretch the levels or use the map to displace geometry. The 16-bit export keeps 65,536 steps taken straight from the model's float output, grayscale only. Blender, Photoshop's 16-bit mode, After Effects and most 3D tools read it directly.

Why does the map look soft, and why cap it at 2048 px?

The network itself only ever sees a 518-pixel view of your photo — that is the input size Depth Anything V2 Small was exported with. Everything above that is upsampled, so a 6000 px depth map carries no more detail than a 2048 px one, it just costs memory. Depth edges are therefore a little soft; if you need them crisp against the original pixels, that is edge-guided upsampling, which is what the Mac app does.

My browser has no WebGPU — can I still use it?

Yes, and it is the same model. Without WebGPU the page loads a 26 MB 4-bit build and runs it on the CPU with WebAssembly, on a single thread. On the machine this page was tested on, a 1280 × 855 photo took 5.5 s that way against 0.32 s on WebGPU. Phones take the CPU path too, which is why the download there is 26 MB rather than 94 MB.

I need normal maps or a real parallax video.

That is the Mac app, AIDepthForge. A browser tab is the wrong place for a minutes-long render: the Mac version does normal maps, batches of photos, full-resolution edge-guided depth and exports actual video rather than a short GIF. This page stays the quick one-photo answer.