Eight petabytes of historical Olympic media sit inside the system the IOC runs with Olympic Broadcasting Services, according to the research. Fox Sports keeps more than 1.9 million archived videos. World Athletics holds more than three million historical results. Every season adds match reports, scores, statistics, athlete profiles, interviews, photographs, broadcasts, commentary and highlights to collections that almost never get pruned. The problem is: Sports media solved storage far more thoroughly than it solved search.
A fan remembers a spectacular comeback but not which season it happened in. A producer remembers what an announcer said during a World Series moment and nothing about the date of the game. An editor needs earlier examples of an unusual performance while the current match is still being played. Searches like those never begin with a filename or a category. They begin with memory and intent.
Why Conventional Search Struggles in Sports
At a high level, the mechanics of search are simple enough. Conventional search matches the words typed into a search box against the words attached to a piece of content. Vector search compares meaning, so it can connect two differently worded descriptions of the same event, action or idea.
Traditional search performs well when the vocabulary is known. A query built around an athlete’s exact name, a team, a date or a competition works perfectly well with keyword and structured search. The trouble starts when the searcher knows what happened but has no idea how the content was labelled.
Sports produces plenty of those situations. A team may be better known by a nickname in one market. An athlete’s name may appear differently across languages. Clubs relocate and rebrand.
The research identifies nicknames, naming variations, historical changes, competition context, natural language statistical questions and event based queries as particularly difficult territory for purely lexical retrieval.
What vector search adds is another way into the collection, one where meaning becomes part of the match. Keywords do not become obsolete in the process. Sports is one of the clearest examples of why both approaches need to coexist: a score demands precision, while a remembered moment demands interpretation. The seven applications below sit mostly on the second side of that problem.
1. Video and Highlight Archive Discovery
Sports video archives grow quickly and lose very little of their potential value. A match broadcast may be most commercially important on the day it airs, but the moments inside it stay useful for years afterward. Records get broken, players return to former teams, anniversaries arrive, rivalries renew, documentaries need archival footage, and social teams suddenly need the perfect ten-second clip from a game nobody has discussed in a decade. The material exists. Finding it is the obstacle.
Conventional video search leans heavily on metadata. A clip labelled with the correct player, game, team and action can be found through those fields. But human memory is rarely that tidy. A producer is more likely to remember “the catch at the sideline in overtime” or “the goal after the goalkeeper came off the line,” and sometimes a commentator’s phrase is the strongest clue available. Semantic retrieval gives those descriptions a chance to work as searches.
Fox Sports offers one of the clearest real-world examples. Its work with Google Cloud allows teams to search more than 1.9 million archived videos using information including broadcast commentary, so a memorable call from a previous game becomes another route back to the footage.
The Olympic Archive Shows the Scale of the Opportunity
The IOC and Olympic Broadcasting Services offer an even larger example. Their Sports AI system manages more than eight petabytes of historical Olympic media. It applies AI to tagging, categorization and multimodal search, with conversational search powered by Alibaba’s Qwen allowing clips to be retrieved through spoken or written requests.
An archive of that size cannot depend on someone knowing exactly how an asset was catalogued decades earlier. Vector and multimodal search make the collection accessible through the content itself, so commentary, descriptions, imagery and contextual information all become paths into material that previously relied on metadata alone. Footage that once sat inert in storage starts working as an active content resource.

2. Conversational Fan Experiences
Sports platforms are usually organized by content type. News sits in one section, scores sit somewhere else, video has its own interface, and statistics may live in an entirely different product. Schedules, standings, player pages and streaming options add more paths again. Fans do not think in those categories. The question is usually something closer to “When do the Leafs play next?” or “What happened the last time these teams met?”
Bringing Different Sources Into One Query
Conversational search can sit above those separate systems and interpret intent first. Some questions need an article, while others need a schedule, a score, a player page or a clip, and the more complicated ones need several sources at once.
Conversational does not mean approximate. If the question is about tonight’s game time, the schedule should provide the answer. If it asks who leads the standings, structured data remains authoritative. Scores, odds, schedules and mutable statistics make poor candidates for answers based on semantic similarity alone.
Vector search helps interpret intent and find relevant context. Exact systems provide exact facts. The separation carries particular weight in sports, where information changes while people are searching for it.
3. Editorial and Broadcast Research
Sports publishing has always required fast research, and live sport compresses the deadline to something close to absurd. A broadcaster notices an unusual statistic. An editor sees a record approaching. A producer wants to know when something similar last happened. The relevant information may be scattered across structured statistics, past articles and years of historical event data, and the search has to understand the question before it can find the evidence. Natural language search makes large sports databases easier to explore without requiring everyone who touches them to know their internal terminology.
FOX Foresight is one example identified in the research, used during the 2025 World Series to support natural language statistical research for broadcasters. The underlying value is easy to understand. An editorial team should not have to translate every interesting question into the internal structure of a statistics database before investigating it, and search can absorb some of that translation work.
None of which turns the AI system into the statistician. The factual answer still needs to come from authoritative data, while the semantic layer makes the information easier to interrogate. In a live environment, shortening the distance between an editorial question and its supporting evidence is worth considerably more than another generic AI writing feature.
4. Historical Context and Milestone Discovery
Sports depends on its own history to an unusual degree. Every record immediately raises another question about the previous record. A milestone produces comparisons with earlier players, a championship revives footage from a previous title, a rivalry game sends editors back through decades of meetings, and a remarkable individual performance creates demand for every similar performance that came before it. The archive grows more valuable when current events can automatically open new paths into older material.
A purely chronological archive makes rediscovery difficult. Older articles and clips slowly disappear beneath newer material unless someone searches for the exact player, date or event.
Semantic relationships provide another route. A platform can connect a new milestone with earlier content describing comparable achievements, similar match situations or related historical narratives. NBA’s AI-supported “Insights” functionality identifies important narratives, performances, and milestones, illustrating the broader demand for systems that can surface relevant context around current events.
5. Multilingual Sports Discovery
Major sports audiences cross borders almost by definition. A single competition may serve fans in dozens of languages. Player names get transliterated differently between scripts. Teams accumulate official names, abbreviations and local nicknames. Commentary describing the same moment varies significantly between markets.
While keyword search reads those differences very literally, semantic retrieval can make them less restrictive. In practice, that might mean connecting a localized query with an English archive, recognizing alternative versions of a team nickname, or retrieving content about the same sporting concept across languages.
Metadata does real work here. Canonical player, team, and event identifiers, for example, provide the strongest link between records, and vector search complements them by improving discovery around them. Translation alone solves nothing if the underlying information architecture treats every variation as an unrelated string. For global sports platforms, the combination is the point.
6. Personalization and Content Routing
Most sports platforms already personalize something. Favourite teams influence feeds, followed competitions affect notifications, viewing history shapes recommendations, and location may determine which events can be watched. Vector representations add another potential signal by capturing relationships among content and interests at a more conceptual level.
Recommendation systems sometimes get less interesting when popularity dominates everything. The largest teams, newest stories, and most-watched highlights keep accumulating attention, while relevant material further down the tail becomes harder to surface. Vector search can support a recommendation engine as an additional relevance signal without becoming one.
The strongest approach still starts with good base relevance. Personalization belongs after the system has identified genuinely relevant material, and the underlying research makes the same point by recommending bounded personalization only once freshness, rights and core relevance have been satisfied.
7. Making Sports Content More Accessible to AI
The final application extends beyond the search box. Organizations are increasingly thinking about how their existing content can support AI assistants, answer experiences, and retrieval-augmented generation, all of which need reliable access to source material. Vector search can provide part of the retrieval layer behind an AI assistant. A question about a historic match could retrieve the relevant article, transcript and event record before a response is generated. A question about a player could pull from approved profile content and current structured information. A search about a remembered highlight could lead directly to the corresponding media asset.
Grounding of that kind separates a useful sports assistant from a model generating an answer out of whatever it happens to remember, and making content understandable and retrievable by approved AI experiences does not require surrendering control over how that content is used. A mature content architecture needs discoverability and governance together, and the tension between them will only grow as AI becomes another interface through which sports audiences encounter information.
The Common Thread Across All Seven
The seven applications share a failure mode. Information that exists inside the organization cannot be found by the people who need it, because it was filed by system logic and is being searched for by human logic. Fields, tags, filenames, databases and site sections on one side; half-remembered moments and vague descriptions on the other.
Keeping Exact Facts in the Right Systems
The strongest sports implementations maintain a clear line between discovery and authority. Semantic retrieval is good at interpreting loosely expressed intent, identifying a remembered moment, connecting related articles and searching descriptions that use different vocabulary. It should never become the source of truth for a live score.
Scores, schedules, odds, standings, and other mutable facts should all be kept in authoritative structured systems. Vector retrieval can determine that a query is asking about a schedule, while the schedule database supplies the answer.
Rights work the same way. Sports content can be restricted by territory, subscription, blackout, language or licensing period, and a highly relevant video result is still the wrong result if the person searching is not entitled to access it. Rights information needs to live in explicit metadata, not in something the semantic layer is expected to infer.
All of which points toward hybrid search as the practical model for sports. Keyword retrieval handles precision, structured systems handle current facts, vector search handles meaning, and metadata establishes context and permissions. Each system keeps the job it performs best.
Building a Smarter Sports Content Layer
Trew Knowledge helps organizations build enterprise digital platforms that make content easier to manage, govern and connect with emerging AI experiences. Sports organizations with deep editorial and media libraries have more at stake than a better search box. Years of accumulated content become usable across fan experiences, editorial workflows and AI-powered discovery. Contact our experts today.
