Product Thumbnail

Jockey by TwelveLabs

The video AI agent that understands your whole library

Productivity
Artificial Intelligence
Video
Visit WebsiteSee on Product Hunt

Hunted byfmerianfmerian

Jockey is the first AI that understands your entire media library just like you do, searching by person, moment, or context across every photo and video you've captured. Powered by TwelveLabs' advanced model stack, Jockey improves automatically with every update. Whether you need to connect via MCP for Claude/ChatGPT or build custom applications using our API, Jockey makes your media library instantly searchable and accessible.

Top comment

I'm Aiden, co-founder and CTO at TwelveLabs. Super excited to share Jockey with you today.


Most media search is still metadata search: filenames, timestamps, maybe some object tags from an older computer vision model. None of that captures what's actually happening in your footage: who's in a scene, what they're doing, the context, dialogue, on-screen text. For that you need models that natively understand time and space in video, not a bag of sampled frames.


That's our stack. Marengo, our embedding model, resolves a query like "the moment we almost missed the flight" to real retrieval across video and images, not keyword matching. Pegasus, our video-language model, segments an entire video on a schema you define and returns structured, timestamped moments. Jockey is a unified agentic system that reasons across your videos and images: a reasoning model plus a memory layer that builds a knowledge store from your corpus, so it can decompose a query, retrieve, segment, and reason across the whole thing.


The point is a corpus-level understanding you can act on. Point Jockey at thousands of videos and images, say "cut me a highlight reel" or "pull the best viral moments," and it comes back with timestamped cuts you can use. The model-only approach can't do that as dumping one video into a context window is bound to a single file, and a single forward pass runs out of room fast. It can tell you about one video; it can't reason across your catalog or build a reel from thousands.


Because the models are what we ship and improve continuously, Jockey's reasoning and retrieval quality improve as we push new versions, meaning no re-integration on your end.


Two ways in:

  • MCP server: connect Jockey as a tool in Claude and query your library directly. ChatGPT coming soon.

  • API: full programmatic access to build custom retrieval or agent workflows on your own library.

This is a research preview, so if you hit edge cases (ambiguous queries, retrieval misses, latency) I want to hear about them. Let us know anytime! 


Best,

Aiden 

Comment highlights

Congrats on the launch. The whole-library search angle is strong, especially for teams with years of product demos, webinars, and raw footage. When Jockey returns a highlight reel or a moment-based answer, how much provenance does the user get back? I would want every suggestion to point to exact source clips and timestamps so the result is easy to verify before editing or publishing.

the split between Marengo for retrieval and Pegasus for schema-based segmentation is a real architecture, not just a wrapper on top of an LLM. one practical question - for personal footage across years of different phones/cameras, is there an upfront processing step you have to run before search works, or does it index incrementally as you add new files to the library?

Was lucky enough to get access to Jockey a week or so ago (thank you team!), and have been blown away at the use cases we've already uncovered. It's changing the way we look at creative strategy across our roster of clients, and we've found some novel ways to extrapolate learnings that are informing some of our performance marketing campaigns that are already showing positive uplift.

Go TwelveLabs!

Finally gave this a spin on a few clips and the semantic search actually nails what's happening in the scene, not just matching keywords. Impressed it picked up on subtle actions without me tagging anything.

A live collaboration mode would be huge, letting a team tag and comment on different timestamps together while the AI pulls those notes into a shared summary report.

Honestly impressed by how well it picks up on subtle visual cues in long videos, not just obvious keywords. Searched a 40 minute documentary for a specific gesture and it nailed it in seconds.

One thing that would make this way more useful for me: a timeline-based search view where I can scrub to the exact moment a concept appears, instead of just getting text hits. Right now it sounds like the results are summaries, but for video editing workflows I really need frame-accurate jumps. Would love to see that built in.

honestly the search across hours of footage feels almost scary good, like you throw in a rough idea and it actually pulls the right moments out without much fuss.

honestly the search is way better than i expected, threw in some random clips and it actually pulled out the exact scenes i was thinking of. pretty wild that you can basically ask it questions about what's happening in the video.

A browser extension that lets you right-click any video and instantly get a chapter breakdown or summary using your Pegasus model would be huge, especially for longer YouTube content or lectures where I don't always want to scrub manually.

finally a video ai that actually finds the moment i describe instead of just dumping timestamps. tried it on some old footage and the semantic search picked out exactly the scene i was thinking of.

Searched a bunch of old vacation clips by typing "sunset over water" and it actually pulled the right moments, which kind of startled me. The natural language search feels way more useful than tagging everything manually.

The way the search results show exact timestamps with preview thumbnails makes me feel like I'm scanning a real video library, not just a text index. That attention to temporal precision shows serious craft.

honestly the search looks solid, but it would be super helpful if you could save and share specific video clips or search results with timestamps baked in. basically a way to send someone straight to the exact moment in the video rather than just linking the whole thing and saying "go to 4:32". that would make it way more useful for team collaboration

A live collaboration mode where teams can annotate and tag specific video segments together in real time would be huge. Right now analysis feels like a solo task, but most of our video review happens in group settings where marketers, editors, and strategists need to align on what they see.

Would love a timeline-based annotation view where I can click any moment in a search result and instantly see the surrounding visual and audio context. Right now results feel like a black box of timestamps, but seeing a quick visual snapshot before clicking through would make reviewing long footage way faster.

The compositional queries are where library-scale video search tends to break. Single-concept stuff like 'red car' works fine off embeddings, but 'the moment right after the door opens' needs temporal grounding that flat similarity search can't reach. When we built multimodal search over video the causal and ordering queries were exactly where recall fell off a cliff. Does Jockey's agent decompose those multi-step queries and reason over ordering, or is retrieval a single embedding lookup under the hood?

honestly the search is way sharper than i expected, like i threw in a random cooking video and it pulled out the exact moment they mentioned "fold in the eggs" without any tagging. pretty cool to actually see video understanding feel useful

would be cool if you added a built-in timeline view so you can jump straight to the exact moment a search result happens in the video. right now i think you only get timestamps in text, which is helpful but still requires manual scrubbing. basically a visual scrubber with the matched segments highlighted would save a lot of clicks.

About Jockey by TwelveLabs on Product Hunt

The video AI agent that understands your whole library

Jockey by TwelveLabs launched on Product Hunt on July 21st, 2026 and earned 210 upvotes and 28 comments, placing #7 on the daily leaderboard. Jockey is the first AI that understands your entire media library just like you do, searching by person, moment, or context across every photo and video you've captured. Powered by TwelveLabs' advanced model stack, Jockey improves automatically with every update. Whether you need to connect via MCP for Claude/ChatGPT or build custom applications using our API, Jockey makes your media library instantly searchable and accessible.

Jockey by TwelveLabs was featured in Productivity (656.7k followers), Artificial Intelligence (474.3k followers) and Video (1.9k followers) on Product Hunt. Together, these topics include over 261.3k products, making this a competitive space to launch in.

Who hunted Jockey by TwelveLabs ?

Jockey by TwelveLabs was hunted by fmerian. A “hunter” on Product Hunt is the community member who submits a product to the platform — uploading the images, the link, and tagging the makers behind it. Hunters typically write the first comment explaining why a product is worth attention, and their followers are notified the moment they post. Around 79% of featured launches on Product Hunt are self-hunted by their makers, but a well-known hunter still acts as a signal of quality to the rest of the community. See the full all-time top hunters leaderboard to discover who is shaping the Product Hunt ecosystem.

Reviews

Jockey by TwelveLabs has received 1 review on Product Hunt with an average rating of 5.00/5. Read all reviews on Product Hunt.

Want to see how Jockey by TwelveLabs stacked up against nearby launches in real time? Check out the live launch dashboard for upvote speed charts, proximity comparisons, and more analytics.