Your search API just got acquired: a diligence checklist after Tavily, Jina and ScrapingBee

Written by Nathan Kessler
Last updated: 8 min read

On 10 February 2026, Nebius announced an agreement to acquire Tavily, one of the most widely used search APIs in agent stacks. The press release does not disclose a price. Bloomberg reported, citing people familiar with the matter, that Nebius would pay $275 million upfront with the figure rising to as much as $400 million on milestones, and that founder and CEO Rotem Weiss would join Nebius with the team. That was the third acquisition in eighteen months to land under an API that AI teams call in production: Elastic bought the company behind Jina AI in October 2025, and Oxylabs' company group bought ScrapingBee in June 2025 for an undisclosed eight-figure sum.
Every piece of coverage I have read on these deals is a recap of the announcement. None of them answers the buyer's question, which is "what changes in my contract, and when do I find out." Acquisition risk in web data is rarely product risk. The product usually keeps working. The risk is that four specific things get rewritten quietly: self-serve pricing, data-retention terms, roadmap priority, and who is on the hook for the SLA.
This is the buyer-side version. It is a different question from who owns the crawl, which I covered in index ownership and the AI search stack. Here the asset changing hands is the company.
Three deals, one shape
The pattern is not consolidation among peers. In all three cases a neutral standalone API became a component inside something larger that sells a different primary product.
Nebius sells GPU capacity and inference. Elastic sells search and observability infrastructure. Oxylabs sells proxies. None of them bought a web data API in order to be in the web data API business. They bought it because it completes a stack they already sell, and that difference is what the questions further down are testing for. A standalone API's incentive is to be excellent at one call and cheap enough that you keep making it. A component's incentive is to make the surrounding platform easier to buy.
That reframing does not automatically hurt you. ScrapingBee is the counterexample worth holding on to: at acquisition the founders stayed, the product stayed a separate entity under its own brand, and the company's announcement told customers there would be no price increase and in some cases lower costs. Founder retention and brand independence are checkable a year later; the pricing promise needs a diff against the published rate card. Acquisition on its own is not a reason to leave a vendor. It does convert a set of terms you had treated as settled into terms somebody else now decides, which is what the checklist below is for.
What the Tavily deal says, and what it does not
Nebius' announcement is unusually specific about continuity and silent about price. It says Tavily keeps its brand and continues serving existing customers, that Weiss leads ongoing product development, and that Tavily gains access to Nebius' infrastructure and engineering resources. Nebius' Roman Chernin framed the combination as letting developers stop managing multiple vendors: Token Factory supplies inference, Tavily supplies real-time web access.
By March the commercial shape was visible. Tavily's own recap of NVIDIA GTC 2026 describes the company as "part of that foundational layer," positions Tavily's search inside Nebius' Token Factory, and sits the whole thing inside what the post describes as NVIDIA's deepening partnership with Nebius, including a $2 billion investment. That is Tavily's own account of its post-acquisition positioning, not a third-party assessment, and it is worth reading exactly as what it is: a vendor describing itself as a layer of somebody else's vertically integrated stack.
Read that with a buyer's eye. "Stop managing multiple vendors" is a benefit if you were already going to buy inference from Nebius. If you are not a Nebius compute customer, the same sentence describes a product whose roadmap is now steered by how well it sells compute.
What a GPU-cloud parent changes structurally
Three mechanisms, all of them ordinary corporate behavior rather than bad faith.
Bundling economics. Once the API is a line in a platform deal, the standalone price stops being the price that matters to the parent. The rational move is to hold or raise standalone list pricing while discounting the bundle, because the bundle carries the compute margin. Standalone buyers do not get a worse product. They get a worse relative deal, gradually.
Roadmap gravity. Feature priority follows the largest revenue line. A search API owned by an inference platform will invest in the things that make agents on that platform work better: latency against co-located models, native integration with the parent's SDK, enterprise controls that show up in procurement. Investment in the things that make the API easy to use from somewhere else is not withdrawn. It just stops being first.
Support and SLA reassignment. This one is mechanical and easy to miss. When an acquired company's operations are folded into the parent, the entity named in your agreement can change, the status page can be consolidated, and the escalation path you tested during evaluation can be replaced by the parent's support tiering, where your spend is small.
The price gap, and what it says about the scarce asset
Elastic's filings put a hard number on the other side of this market. The Q3 fiscal 2026 filing records the acquisition of Conic AI Technology Limited, the entity behind Jina AI, on 7 October 2025, for total purchase consideration of $43.3 million, of which $6.9 million is held back for indemnity obligations and released on the 24-month anniversary. The allocation is the interesting part: $6.5 million to developed technology, amortized straight-line over a two-year useful life, and $32.3 million to goodwill, which Elastic attributes to enhancing its Search AI solutions and to the acquired workforce.
Set that against the reported Tavily figure and the comparison is rough but directional. Roughly $275 million reported for live web access against $43.3 million filed for an embeddings, reranker and reader business. The two numbers are not strictly comparable: one is journalism about a deal with milestone components, the other is an audited line in an SEC filing. But the gap is large enough to survive the caveats, and Elastic's own two-year amortization on the developed technology says something blunt about how durable a buyer considers model-side IP to be.
The market is pricing live, permissioned access to the current web well above the models that process it. If you are choosing a vendor, that is the asset to check ownership of, and it is why pricing pressure in this category keeps landing on access rather than on inference.
The term that gets rewritten first
Data-retention and processing terms are usually the earliest post-close change, and they are the terms your compliance review already signed off on.
The reason is structural. A parent with an existing enterprise security posture cannot run two data-processing regimes, so the acquired product's DPA, subprocessor list, retention window and regional storage commitments get harmonized onto the parent's. That process often improves the terms. It also invalidates the specific document your legal team reviewed, and nobody sends you a diff.
If your product handles anything covered by a customer commitment, put a calendar reminder on the acquired vendor's terms pages and diff them. The subprocessor list is the one to watch: a new parent adds itself and its infrastructure providers, and if your own customer contracts require subprocessor notification, that change propagates into your obligations.
The eight questions
Ask these of any vendor in the twelve months after an acquisition. In writing, to a named person.
- Is self-serve still a supported product line, or a funnel into sales? Ask whether a customer can stay on self-serve indefinitely at growing volume without a mandatory sales conversation.
- What is the notice period on tier and pricing changes? Get a number of days. If the answer is "we would communicate in advance," that is not a term.
- Do the retention and processing terms survive the transition, and when does the current DPA expire? Ask for the subprocessor notification commitment specifically.
- Which legal entity owns the SLA now, and did the remedies change? Uptime credits written by a startup and uptime credits written by a public parent are rarely the same document.
- Is there a bundling requirement or bundle-only pricing coming? Ask whether any current or planned feature is gated to customers of the parent's other products.
- What happened to the public roadmap and changelog cadence? Compare shipping frequency for the six months before close to the six after.
- Are the founders and the core engineers still there, and through what date? Retention packages have end dates. Ask when.
- What is my exit ramp? How I export what I have, how far the response schema is from an alternative's, and how long a swap takes.
Question eight is the only one whose answer you fully control.
Reading the signals before the announcement
The independence markers worth tracking are boring and public. Docs moving onto the parent's domain. The status page folding into the parent's. A pricing page edit that removes an annual option or repositions the free tier as a trial. Terms and DPA revision dates changing without an email. Support moving from a shared channel to a ticket portal. Job postings for the acquired product listing the parent's team names. None of these is decisive alone. Three of them in a quarter is a roadmap you were not shown.
Portability is the hedge
You cannot underwrite a vendor's ownership. You can make replacement cheap, and that is the only durable protection. Write your retrieval layer against an internal interface rather than a vendor SDK, normalize responses to your own shape at the boundary, and keep a second provider wired and periodically exercised rather than merely documented.
For search, the alternatives that keep a swap cheap are the ones with comparable single-call semantics: Exa, Parallel AI, Linkup, Brave Search API, Valyu and Ceramic. They differ on index ownership, licensing posture and result shape, which is exactly the axis to test on before you need to. Our guide to choosing a search API walks the trade-offs. For URL-to-markdown, Jina AI and Firecrawl return close enough output that a working adapter between them is a day of work, not a migration.
What this means if you build on web data
Acquisition risk is contract risk, not quality risk. The product usually keeps working. Pricing terms, retention terms and SLA ownership are what move, and they move without an announcement.
A component's incentives are not a standalone's. When the parent sells compute or infrastructure, the API's job becomes making that easier to buy. Budget for the standalone tier becoming the worse relative deal even if the list price never changes.
Diff the terms pages on a schedule. Data-processing and subprocessor terms are the first thing harmonized onto a parent's standard, and your compliance review approved the old version.
The exit ramp is the only lever you own. Every hour spent on a provider-neutral retrieval interface is insurance against a deal you will read about the same morning everyone else does.
- #market-analysis
- #api-selection
- #ai-search-apis
- #pricing
More from the blog
- Exa and Parallel are priced like infrastructure now. Price your dependency accordingly
Jul 27, 2026 · 7 min read
- One agent task, twenty search calls: a cost model for agentic search pricing
Jul 27, 2026 · 8 min read
- Crawler purpose is now a product spec: reading Cloudflare's September 15 default as a buyer
Jul 27, 2026 · 8 min read