The final step in semantic search: Is Rerank effective?
Community Discussion · Tracks

The final step in semantic search: Is Rerank effective?

Pao Tiao XianPao Tiao XianSep 212026/09/21 158 views

A friend recommended Rerank to me, so I tried it out to see how useful it actually is. He showed me a screen recording where searching for which receipts employees need to upload for travel reimbursement resulted in keyword searches mixing invoice titles with subsidy standards, pushing the relevant documents down to fourth or fifth place. He said this job belongs to reranking models—Rerank, literally meaning "re-sort." The method involves first roughly recalling a batch of candidates, then comparing and scoring them against the query one by one. It doesn't generate answers; it just puts results in the right order.

I've been tinkering with a search agent recently, with Feishu-exported docs, ticket descriptions in Postgres, and a small amount of semi-structured data at hand. Previously, vector retrieval would recall dozens of candidates, pass them to an LLM for answering, and the results were hit-or-miss. I tried integrating Cohere Rerank into the retrieval pipeline. After getting the API key, I wrote a small script taking query and documents as input and outputting relevance score. The interface isn't flashy; it's mostly JSON. The tricky part is chunking: chunks too long eat costs, chunks too short break context. Initially, I only passed titles, and invoice titles ranked very low; adding body text fixed the ordering.

There were many bottlenecks. At first, I followed standard API habits and passed text, but the field name was wrong, and error messages weren't friendly—I spent half a day checking documentation. Later, switching to rerank parameters common in search services like Elasticsearch made the calls stable. Cost is also an issue. According to the pricing info I saw, Cohere charges $2 per thousand searches, which differs from the common token-based billing. This metric suits stuffing dozens of long documents at once, but might not be cost-effective for high-frequency queries on short texts. My testing suggests you shouldn't infer business costs from cheap demos; estimate using real logs first.

The second run had some surprises. For the same batch of tickets, dozens of candidates were recalled, but I let Rerank return only the top ten. Items previously ranked lower regarding overseas exchange rates and receipt requirements rose to the top because they covered location, receipts, and deadlines simultaneously. Rough recall is like casting a net; Rerank is like picking fish by size. I also tried mixed Chinese-English queries, and the model wasn't completely misled by keywords. When semi-structured tables retained headers, it better understood the relationship between fields and values.

You can't just look at effectiveness. I picked dozens of common questions, manually labeled ideal answers, and checked where the first relevant document ranked—similar to ranking metrics like MRR. Before adding reranking, several answers were pushed to the second screen; after adding it, most questions had hits in the top two. But not all documents improved. Clickbait titles, duplicate policies, and expired notices got sorted more neatly by Rerank—neatly wrong. It doesn't judge facts for you. I previously wrote about breaking AI compute news into fact lists; the more automated retrieval becomes, the more manual verification of scope and exclusions is needed.

It's worth integrating if you already have vector databases, enterprise search, or customer service knowledge bases. It's like fine-ranking in recommendation systems: coarse ranking ensures nothing is missed, while fine ranking puts the most clickable items upfront. If documents aren't organized, there are no query logs, latency budget is only tens of milliseconds, or you just want a pretty demo, you'll likely get criticized. My single rerank operations mostly complete within hundreds of milliseconds. Documentation mentions that even with many candidates, the Cohere API takes about six hundred milliseconds, which is acceptable, but you can't stack too many calls.

Rerank isn't a silver bullet, but it can clean up the messiest sorting parts of search. Next, I plan to insert it between WorkBuddy and my search agent for A/B testing, comparing recall-only vs. recall-plus-rerank, looking at click-through rates and manual correction rates. If multilingual content, long documents, and semi-structured data increase, vector similarity alone may not suffice. Whether independent rerank APIs can continue to sell depends on whether they are cheaper, faster, and more controllable than built-in model ranking.

1 replies

?
Ctrl + Enter to reply
Tang
TangSep 23

At two dollars per thousand calls, my tiny query volume on the night shift makes it completely unaffordable.