LocateAnything Visual Grounding
An open-vocabulary visual grounding solution via the LocateAnything extension β based on the LocateAnything-3B vision-language model, it finds / counts arbitrary objects in a frame by natural language, with phrase grounding, text localization, GUI grounding, and pointing. Supports zero-shot detection and counting for classes a fixed detector doesn't know or ad-hoc needs.