Skip to content

OmniJev/OneJev

114 12PythonApache-2.0Updated 3 days ago
View on GitHub

🚀🚀 A multimodal System One decision model that gives calibrated answers to typed questions about screens, photos, video and text in one forward pass.

calibrationcomputer-usedecision-modelgui-agenthuggingfacejevllmmultimodalpytorchqwensystem-onetypesafevideo-understandingvision-language-modelvlm

Project homepage

Our verdict: worth keeping an eye on

multimodal decision model that requires large pretrained models and heavy dependencies, useful later but not a small installable block now

Filed under browser automation in our directory.

Similar repos