TwelveLabs, a San Francisco and Seoul-based company building AI technology for video understanding, announced on July 1, 2026 that it has raised $100 million in a Series B funding round to build what it calls "video superintelligence." The round was co-led by NEA and NAVER Ventures, with participation from Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital and Red Bull Ventures.
The company's platform enables machines to understand video content at a semantic level, opening applications in search, content moderation, editing, and analysis that were previously impractical. As part of the round, Amazon, a repeat investor, secured AWS as TwelveLabs' preferred cloud, with new models tuned for AWS Trainium chips launching there first.
Last updated August 4, 2026: added real deployed customer use cases (auto-repair triage, legal discovery, insurance claims), the Rodeo product, and honest competitive context against Google's Gemini, none of which were in the original funding-announcement recap.
What Happened: TwelveLabs' $100M Series B
TwelveLabs announced its Series B funding, bringing total investment to over $150 million. The company builds Marengo, its perception model, and Pegasus, its reasoning model, which together index and query existing video footage across up to 2 hours of context, forming what TwelveLabs describes as a full-stack agentic intelligence system for video that combines perception, knowledge and reasoning, distinct from video generation tools.
The investor list includes strategic participants like Amazon and NAVER, and the round comes with a concrete commercial arrangement: Amazon's investment secured AWS as TwelveLabs' preferred cloud provider, with new TwelveLabs models tuned for AWS Trainium chips launching there first. TwelveLabs plans to use the funding to accelerate R&D and open new offices in New York and London alongside its existing San Francisco and Seoul operations.
Key Details
TwelveLabs' technology goes beyond simple object detection or transcription. Its models understand the semantic content of video, enabling queries like "find scenes where someone explains a concept" or "identify moments of conflict in this footage."
The platform supports multiple languages and can work with video content at scale, indexing and querying footage across up to 2 hours of context per video. Applications include media asset management, automated content tagging, video search, and real-time analysis for live streams.
Who's Actually Using This
TwelveLabs says more than 30,000 developers use its platform, and the real-world applications go well beyond generic "video search" pitches: a large auto manufacturer built an automated system that reviews footage of cars arriving for service and provides initial mechanical analysis; law firms use it to work through video-heavy discovery and investigative evidence; insurance companies are converting paper-based claims processes into video-based ones that need less manual review; and municipalities use it for real-time threat detection and traffic management. The company also recently released Rodeo, an AI copilot that lets video editors search and assemble footage using natural-language prompts, a more concrete product than "video superintelligence" branding alone conveys.
Why TwelveLabs Instead of Gemini or Copilot's Video Features
The honest competitive picture is that Google's Gemini can already search through footage, and Microsoft and Amazon both offer video analytics services for spotting objects in clips, none of this is a market TwelveLabs invented from nothing. TwelveLabs' pitch is that it was built video-first rather than as a feature bolted onto a general-purpose multimodal model, and that its models can be customized on a customer's own video data in ways general models aren't set up for. That differentiation matters because Google and Meta are pouring consumer-advertising-scale resources into their own in-house video understanding, real, well-funded competition, not a hypothetical one.
The split between Marengo (perception) and Pegasus (reasoning) reflects a deliberate architectural choice rather than an arbitrary product naming decision. Marengo's job is converting raw video, pixels changing over time, into a structured representation the system can search and query: identifying objects, actions, scene changes, and spoken or on-screen text across the footage. Pegasus then reasons over that structured representation to answer more complex, open-ended questions, like "find moments of conflict" or "summarize what happened in this footage," that require connecting multiple detected elements into a coherent interpretation rather than just returning a list of detected objects. Separating perception from reasoning into two purpose-built models, rather than asking a single general-purpose multimodal model to do both, is TwelveLabs' technical bet: that specialization at each stage produces better video understanding than treating video as just another input type alongside text and images inside a broader model, the approach Gemini and similar general models take.
Why It Matters
Specific, already-deployed use cases like auto-repair triage and insurance-claims automation are a stronger signal than the funding total: they show real organizations paying for this today, not betting on a future capability. The open question is whether video-first specialization stays defensible as Gemini and other general models keep improving at video understanding as a side effect of broader multimodal training.
What Happens Next
TwelveLabs will use the funding to expand model capabilities, grow engineering, and open New York and London offices alongside its San Francisco and Seoul teams. Whether specialized video models stay ahead of general-purpose multimodal models on video-specific tasks, or get absorbed as those models improve, is the competitive question that will determine if this round was well-timed or premature.
Final Takeaway
TwelveLabs' round is backed by concrete deployed use cases, not just a large addressable market pitch, and the AWS Trainium partnership gives it real infrastructure backing. Its real test is staying differentiated from Gemini and other general-purpose models that are getting better at video understanding as a byproduct of broader development, not as their core focus.
Key Points
- TwelveLabs reports more than 30,000 developers on its platform, with deployed use cases including automated vehicle-inspection triage, legal discovery review, and insurance claims processing.
- Its competitive pitch against Google's Gemini and Microsoft/Amazon's video analytics tools rests on being video-first and customizable on customer data, not a bolt-on feature of a general-purpose model.
- TwelveLabs announced its Series B funding on July 1, 2026, bringing total investment to over $150 million.
Video Understanding Applications
TwelveLabs' technology enables applications that were previously impractical due to the difficulty of analyzing video content at scale. Media companies can search archives of footage using natural language queries rather than manual tagging. Security organizations can analyze surveillance footage for specific events or behaviors. Educational platforms can index video content for searchable learning materials.
The company's multimodal approach, analyzing visual, audio, and textual elements together, produces more accurate understanding than single-modality analysis. A scene might be understood differently based on visual content, spoken dialogue, and on-screen text. Integrating these signals produces richer semantic understanding.
Competition in video understanding includes general-purpose multimodal models from Google, OpenAI, and Meta. TwelveLabs differentiates through specialization, arguing that dedicated video models outperform general models for video-centric applications.
FAQs
Sources and Verification
- GlobeNewswire (official release), July 1, 2026
- Tech Funding News, July 2026
- Silicon Report: competitive positioning against Google
This article was reviewed as part of CapisTech's editorial fact-checking process.



