Terrier is billed as a highly flexible, efficient, and effective open source search engine, readily deployable on large-scale collections of documents. Terrier implements state-of-the-art indexing and retrieval functionalities, and provides an ideal platform for the rapid development and evaluation of large-scale retrieval applications.
Terrier follows a plugin architecture, and is easy to extend to develop new retrieval techniques, add new ranking features or experiment with low-level functionality such as index compression.
It’s written in the Java programming language, and therefore runs on all main operating systems.
Features include:
- Indexing support for common desktop file formats, and for commonly used TREC research collections (e.g. TREC CDs 1-5, WT2G, WT10G, GOV, GOV2, Blogs06, Blog08, ClueWeb09, ClueWeb12).
- Many document weighting models, such as many parameter-free Divergence from
- Randomness weighting models, Okapi BM25 and language modelling.
- Supervised (machine learned) ranking models are supported via learning to rank.
- Conventional query language supported, including phrases, and terms occurring in tags.
- Handling full-text indexing of large-scale document collections, in a centralised architecture to at least 50 million documents, and using the Hadoop MapReduce distributed indexing scheme for even larger collections.
- Incremental indexing and retrieval capabilities to support real-time search
- Modular and open indexing and querying APIs, to allow easy extension for your own applications and research.
- Active Information Retrieval research fed into the Open Source platform.
- Indexing:
- Out-of-the box indexing of tagged document collections, such as the TREC test collections.
- Out-of-the box indexing for documents of various formats, such as HTML, PDF, or Microsoft Word, Excel and PowerPoint files.
- Out-of-the box support for distributed indexing in a Hadoop MapReduce setting.
- Indexing of field information, such as the frequency of a term in a TITLE or H1 HTML tag.
- Indexing of position information on a word, or a block (e.g. a window of terms within a distance) level.
- Support for various encodings of documents (UTF), to facilitate multi-lingual retrieval.
- Support for changing the tokenisation being used.
- Updatable indices to support real-time search
- Indexing support for query-biased summarisation.
- Support for fetching files to index by HTTP, allowing intranets to be easily searched.
- Highly compressed index disk data structures with built-in pluggable compression algorithms.
- Highly compressed direct file for efficient query expansion.
- Alternative faster single-pass and MapReduce based indexing.
- Various stemming techniques supported, including the Snowball stemmer for European languages.
- Retrieval:
- Provides desktop, command-line and Web based querying interfaces.
- Provides standard querying facilities, as well as Query Expansion (pseudo-relevance feedback).
- Can be applied in interactive applications, such as the included Desktop Search, or in a batch setting for research and experimentation.
- Provides many standard document weighting models, including up to 126 Divergence From Randomness (DFR) document ranking models, and other models such as Okapi BM25, language modelling and TF-IDF. Two new 2nd generation DFR weighting model, JsKLs and XSqrA_M, are also included, which provide robust performance on a range of test collections without the need for any parameter tuning or training.
- Advanced query language that supports synonyms, +/- operators, phrase and proximity search, and fields.
- Learning-to-rank support enables out-of-the-box supervised ranking models.
- Provides a number of parameter-free DFR term weighting models for automatic query expansion, in addition to Rocchio’s query expansion.
- Flexible processing of terms through a pipeline of components, such as stopword removers and stemmers.
Website: terrier.org
Support: Documentation, GitHub Code Repository
Developer: School of Computing Science, University of Glasgow
License: Mozilla Public Licence
Learn Java with our recommended free books and free tutorials.
Related Software
| Desktop Search Engines | |
|---|---|
| Recoll | Desktop search tool with full text search. Based on the Xapian search engine |
| TinySPARQL | Fiesystem indexer, metadata storage system and search tool |
| Baloo | KDE's file indexing and search solution |
| Terrier | Flexible, efficient, and effective search engine |
| DocFetcher | Fast document search with this desktop search software |
| File Brain | Intelligent desktop file search application |
| dsearch | Filesystem search service |
| Searchmonkey | Java desktop search engine |
| Pinot | Personal search and metasearch |
Read our verdict in the software roundup.
Explore our comprehensive directory of recommended free and open source software. Our carefully curated collection spans every major software category.This directory is part of our ongoing series of informative articles for Linux enthusiasts. It features hundreds of detailed reviews, along with open source alternatives to proprietary solutions from major corporations such as Google, Microsoft, Apple, Adobe, IBM, Cisco, Oracle, and Autodesk. You’ll also find interesting projects to try, hardware coverage, free programming books and tutorials, and much more. Discovered a useful open source Linux program that we haven’t covered yet? Let us know by completing this form. |


Please read our Comment Policy before commenting.