No description
Find a file
Prashant Pandey 4515d34e8d paper
2026-08-30 17:28:12 +05:30
assets add image 2026-08-30 11:37:58 +05:30
src facebook/dinov2-giant 2026-08-30 11:29:52 +05:30
.gitignore hide cache 2026-08-29 18:45:06 +05:30
pyproject.toml readme path 2026-08-30 11:44:59 +05:30
readme.md paper 2026-08-30 17:28:12 +05:30
uv.lock lockfile 2026-08-29 18:47:12 +05:30

visual embeddings of multiples images in a scatter plot. uses facebookresearch/dinov2 for embeddings.

the input images are passed through the vision transformer to get semantic feature vectors(and are cached).

the embeddings are projected down to 2d coordinates using umap/tsne/pca. (choice can be done at the webapp as per number of images).

on the very first run, model weights are downloaded and dumped to cache. (current base model is ~4.5gbs, you can change it as per the comments on embed.py)

visualization of embeddings:

usage

git clone https://github.com/pandey-ps/space.git
cd space
uv run python src/app.py