7 ms·
Anyone using vision to parse screenshots? QVQ was too slow. Will give this a shot.
by adamsiem 1y ago
Anyone using vision to parse screenshots? QVQ was too slow. Will give this a shot.
- abrichr 1y agoYou might be interested in https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt
- logankeenan 1y agoI used molmo to parse screenshots in order to detect locations of UI elements. See the repo below. I think Omni parser from Microsoft would also work well. https://github.com/logankeenan/george https://github.com/logankeenan/george https://github.com/microsoft/OmniParser https://github.com/microsoft/OmniParser