NVIDIA has released another open-source 3B model — LocateAnything-3B. One sentence sums it up: you say a phrase, and it precisely boxes the target in an image or video.
The task coverage is unusually broad:
| Capability | Plain-language explanation | |---|---| | Object detection | Class name/description in → boxes out | | Phrase grounding | Natural-language phrase → box the corresponding object | | OCR / text localization | All text in the image + positions | | Document layout analysis | Paragraphs / headings / table regions | | GUI element grounding | Button description → box the UI control | | Point localization | Description → point coordinates |
The technical core is Parallel Box Decoding — producing multiple boxes in one forward pass, noticeably faster than traditional auto-regressive detection approaches. Training data is NVIDIA's own LocateAnything-Data: 12M images spanning two domains, "Enterprise Intelligence" and "Physical AI."
Sources:
- Official page: https://research.nvidia.com/labs/lpr/locate-anything
- Paper PDF: https://research.nvidia.com/labs/lpr/locate-anything/LocateAnything.pdf
- Hugging Face: https://huggingface.co/nvidia/LocateAnything-3B (released 2026-05-26)
- GitHub: https://github.com/gammahazard/locate-anything (runs with a simple
docker compose up)
A few points worth highlighting:
1. In essence, it swaps "specialist models" like YOLO for a "unified VLM" YOLO is absurdly fast on the narrow path of pure object detection (YOLOv10 / RT-DETR are already deployed at industrial scale), and LocateAnything-3B can't beat specialist models on single-task throughput. Its selling point is one model doing six jobs — dropping deployment cost and engineering complexity by an order of magnitude. This is the vision-world version of "a general LLM kills a pile of fine-tuned task models," as happened in the RAG era.
2. "GUI grounding" is a seriously underrated capability Give it "that blue submit button" and the model boxes it — exactly the underlying primitive for GUI agents like Anthropic Computer Use and OpenAI Operator. By training it in the same framework as object detection, NVIDIA is betting that "the agent visual stack will also move toward unification."
3. NVIDIA's open-source strategy still sells GPUs Same playbook as Nemotron and Llama-Nemotron: the model is free, but to run it fast and reliably → buy NVIDIA hardware. "Runnable on consumer GPUs" + "a single docker compose up" lowers the barrier to trying it, but production deployments will still end up on H100/H200. Open source is a business, not charity.
4. An overlooked detail: Parallel Box Decoding Traditional grounding models emit boxes one by one, auto-regressively; parallel decoding emits N boxes at once. In video-streaming scenarios (where every frame needs boxes), the throughput advantage is amplified. My bet: a "LocateAnything-Video" variant appears within a year.
---
Open question:** Do you think "unified VLMs replacing specialist vision models" will actually pan out? Or will "small-and-focused" models like YOLO keep winning industrial deployments — much like RAG, where the theory isn't elegant but the engineering just works?