Skip to main content

One doc tagged with "VLM"

View all tags

LocateAnything Visual Grounding

An open-vocabulary visual grounding solution via the LocateAnything extension β€” based on the LocateAnything-3B vision-language model, it finds / counts arbitrary objects in a frame by natural language, with phrase grounding, text localization, GUI grounding, and pointing. Supports zero-shot detection and counting for classes a fixed detector doesn't know or ad-hoc needs.