DoTA-RAGDynamic of Thought Aggregation RAG

SIGIR LiveRAG 2025 — Oral presentation

Saksorn Ruangtanusak1, Natthapath Rungseesiripak1, Peerawat Rojratchadakorn1, Monthol Charattrakool1, Natapong Nitarach2

1 SCBX2 SCB 10XBangkok, Thailand

#1 industry team · 5th overallSIGIR 2025 LiveRAG ChallengeExplore the standings ↓
DoTA-RAG pipeline: rewrite the query, classify namespaces with four routing samples, embed the query with Snowflake Arctic, retrieve candidates through dynamic routing, prune and rerank, then generate an answer with Falcon3-10B.
Figure 2 · The DoTA-RAG pipeline. Query rewriting, dynamic routing, and hybrid retrieval connect a question to relevant evidence before answer generation. Original artwork from the paper. View full size.

Search the right part of the web

DoTA-RAG combines topic-aware routing with a sequence of retrieval and ranking steps to answer questions over FineWeb-10BT. It targets a practical challenge: finding useful evidence in a large, varied corpus while keeping response time manageable.

Developed for the SIGIR 2025 LiveRAG Challenge, the system uses Falcon3-10B-Instruct for generation and Snowflake Arctic embeddings for retrieval. Its evaluation pairs a diverse internal benchmark with a separate live challenge test.

15Mdocuments in the corpus

500internal benchmark questions

35.63 sper question, final internal configuration

Route first. Retrieve broadly. Refine the evidence.

  1. Rewrite the question. A low-temperature rewrite corrects noisy wording and misspellings before retrieval.
  2. Choose the relevant namespaces. Falcon3-10B-Instruct makes four independent topic classifications. Self-consistency voting selects the top two Pinecone namespaces, which are queried in parallel.
  3. Retrieve and rank. Arctic-embed-m-v2.0 retrieves 100 candidate passages. BM25 narrows these to 20; Cohere Rerank 3.5 selects the final 10.
  4. Assemble the context. The passages are combined within an 8,000-token budget, with proportional truncation when needed.
  5. Generate an answer. Falcon3-10B-Instruct answers using the retrieved context and rewritten question.

The paper reports a 92% smaller average search space with routing. In the internal ablation, time per question falls from 100.84 seconds with Arctic-M to 19.01 seconds after routing is added.

Source: paper §2 and Table 3. The final pipeline adds pruning, reranking, and rewriting after the routing-only configuration.

One corpus, many kinds of knowledge

WebOrganizer classifies FineWeb-10BT along two axes: 24 topics and 24 document formats. Topic labels define the namespaces used for routing; the topic–format distribution also informs the internal benchmark’s sampling.

Figure 6: FineWeb-10BT treemaps. Topics on the left use purple; document formats on the right use orange. Finance and Business is the largest topic at 9.11 percent. News Article is the largest format at 17.40 percent, followed by Personal Blog at 15.64 percent and Product Page at 14.85 percent. Block area represents document count.
Figure 6 · Topics and document formats in FineWeb-10BT. Block areas represent the number of documents in each category, based on WebOrganizer classifiers. Both original treemaps and their colors are preserved. Scroll sideways on small screens or open the full-size figure. Source: paper, Appendix C.

A benchmark built for variety

The 500-question MorganaMultiDocQA benchmark is constructed with DataMorgana and stratified topic–format sampling. It includes single-document and multi-document questions, spanning comparisons, temporal change, procedures, causal explanations, quantities, and verification.

First among industry teams. Fifth overall.

Ped100X / DoTA-RAG brings together researchers from SCBX and SCB 10X. It is the highest-ranked industry team in the supplied standings; the four teams above it represent academic or research institutions.

The industry distinction is also reported in SCBX’s announcement. It describes the team’s placement, rather than a separate award category.

SIGIR 2025 LiveRAG competition standings · 25 entries in the supplied results snapshot. Rank follows the source; scores are rounded to three decimals.
RankTeamInstitutionBordaCorrectnessFaithfulness
1RMIT-ADMSRMIT, Australia7.7071.1990.477
2RAGtifierL3S Research Center, Leibniz University Hannover, Germany7.3511.1340.552
3UDInfoUniversity of Delaware, USA7.2401.2010.623
4MagikarpInstitute of Automation, Chinese Academy of Sciences, China7.0771.2320.656
5Ped100X · DoTA-RAGOur team · #1 industry teamSCBX, Thailand6.2260.9290.043
6ScaledRAGUniversity of Massachusetts Amherst, USA6.0720.9960.418
7HLTCOEJohns Hopkins University, USA6.0191.0700.341
8RagmatazzOpenSource Connections, Germany5.4711.0120.519
9PRMAS-DRCAIISER Kolkata, India5.2070.9230.411
10Hybrid Search with GraphSouthwest University, China5.0770.8750.316
11RUC DeepSearchRenmin University of China5.0670.9690.388
12Graph-Enhanced RAGHuawei Technologies, United Kingdom4.8030.8760.529
13EmoragEmory University, USA4.6830.8910.557
UIUC-RAGentsUniversity of Illinois Urbana-Champaign, USA0.565-0.303
UiS-IAIUniversity of Stavanger, Norway0.5520.434
RAGentATU Dresden, Germany0.8360.200
StarlightCarnegie Mellon University, USA0.8180.433
BagBagHefei University of Technology, China0.694-0.911
UniClustRAGAthens University of Economics and Business, Greece0.6850.460
METURAG0.6730.325
DeepRAGNew York University, United Arab Emirates0.5660.098
SNU-LDILabSeoul National University, South Korea0.5170.103
Gravitational LensUniversity of Auckland, New Zealand0.377-0.988
NoobRAGTU Dresden, Germany0.6550.155
AugmentRAG-TUDTU Dresden, Germany0.5330.656

An em dash means no rank or value is provided. Unranked entries retain their source order. These competition results are separate from the internal evaluation below.

Download the complete source CSV · Includes all metrics, affiliations, paper links, and full-precision scores.

Internal evaluation: each stage has a tradeoff

The ablation below uses the internal test set and Claude 3.5 Sonnet as judge. Correctness ranges from −1 to 2; faithfulness ranges from −1 to 1. These are judge scores, not accuracy percentages.

Table 3 · Full-text scores (“All words”). Each row adds to the preceding configuration. Higher scores are better; lower time is faster.
ConfigurationCorrectnessFaithfulnessSeconds / question
Baseline0.752−0.496
+ Arctic-M1.616−0.216100.84
+ Routing1.562−0.10819.01
+ Pruning1.5620.42829.84
+ Rerank1.6520.67235.20
+ Rewrite · DoTA-RAG1.4780.64035.63

Baseline timing is not reported because its pre-built index is not directly comparable to the authors’ FineWeb-based Pinecone index. Bold values identify the best reported value in each column.

Reranking gives the strongest internal scores. Adding rewriting lowers those scores; the authors retain it to address noisy queries encountered during the live challenge.

Live Challenge Day: a separate evaluation

The paper reports 0.929 correctness and 0.043 faithfulness for the final pipeline on Live Challenge Day. These results come from a different evaluation and should not be compared as another row of the internal ablation.

The authors attribute the low live faithfulness score to an overlooked 300-word output limit. Their subsequent re-evaluation with a tailored Claude judge found faithfulness of 0.702 for uncut answers and 0.336 for capped answers; these diagnostic scores are separate from the official live scores.

Source: paper §4, Table 3.

Featured in

SCBX’s DoTA-RAG places fifth overall and leads industry teams at SIGIR 2025 LiveRAG.

News, press coverage, and social posts about the team’s result, including republications of the SCBX announcement.

Entries marked “Archive” link to the supplied publication listings.

Citation

@misc{ruangtanusak2025dotarag,
  title = {DoTA-RAG: Dynamic of Thought Aggregation RAG},
  author = {Saksorn Ruangtanusak and Natthapath Rungseesiripak and
            Peerawat Rojratchadakorn and Monthol Charattrakool and
            Natapong Nitarach},
  year = {2025},
  eprint = {2506.12571},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  doi = {10.48550/arXiv.2506.12571},
  url = {https://arxiv.org/abs/2506.12571}
}