Search systems enable fast, flexible queries over large datasets. When your database's LIKE queries aren't enough, dedicated search infrastructure becomes essential.
Why Dedicated Search Systems
Database limitations:
- Full-text search is slow at scale
- Complex queries are expensive
- Fuzzy matching, typo tolerance, synonyms are hard
- Ranking/relevance requires custom logic
Search systems provide:
- Inverted indexes for fast full-text search
- Built-in relevance ranking
- Faceted search and aggregations
- Typo tolerance and fuzzy matching
- Horizontal scaling for large datasets
If users need to search through text content—products, articles, users, logs—you likely need a search system.
Search Technologies
| Popular Search Solutions | |
|---|---|
| Name | Description |
Elasticsearch | Industry standard. Powerful, flexible, great ecosystem. Complex to operate. Best for: general search, logging, analytics. |
OpenSearch | AWS fork of Elasticsearch. Similar features, different licensing. Best for: AWS-native applications. |
Algolia | Managed search-as-a-service. Easy to use, fast. Higher cost at scale. Best for: e-commerce, simple use cases. |
Meilisearch | Open source, easy to set up. Good for smaller scale. Best for: startups, simpler needs. |
Typesense | Open source, typo-tolerant. Fast, easy. Best for: applications needing typo tolerance. |
Solr | Apache project, mature. More complex than Elasticsearch. Best for: legacy systems, specific enterprise needs. |
Core Concepts
Inverted Index
The key data structure for fast search.
Traditional index:
Document 1: "The quick brown fox"
Document 2: "The lazy dog"
Inverted index:
"the" → [Doc 1, Doc 2]
"quick" → [Doc 1]
"brown" → [Doc 1]
"fox" → [Doc 1]
"lazy" → [Doc 2]
"dog" → [Doc 2]
Searching "quick fox" → intersect posting lists → Document 1
Documents and Fields
{
"index": "products",
"document": {
"id": "123",
"title": "Wireless Bluetooth Headphones",
"description": "High-quality wireless headphones with noise cancellation",
"category": "electronics",
"price": 149.99,
"rating": 4.5,
"tags": ["wireless", "bluetooth", "noise-cancelling"]
}
}
Analyzers
Transform text for indexing and searching.
Input: "Running Shoes for Women"
Analyzer steps:
1. Lowercase: "running shoes for women"
2. Tokenize: ["running", "shoes", "for", "women"]
3. Stop words: ["running", "shoes", "women"] (removed "for")
4. Stemming: ["run", "shoe", "woman"]
Final tokens indexed: ["run", "shoe", "woman"]
| Common Analyzer Components | |
|---|---|
| Name | Description |
Tokenizer | Splits text into tokens. Standard (word boundaries), whitespace, n-gram. |
Lowercase filter | Converts to lowercase for case-insensitive search. |
Stop words filter | Removes common words (the, a, is) that don't add meaning. |
Stemmer | Reduces words to root form (running → run). |
Synonym filter | Expands queries with synonyms (couch → sofa). |
Search Architecture
Basic Pattern
Application → Search Service → Elasticsearch
↘ Primary Database (source of truth)
- Write to primary database
- Sync changes to search index (async or sync)
- Search queries go to Elasticsearch
- Details fetched from primary database if needed
Keeping Search in Sync
| Sync Strategies | |
|---|---|
| Name | Description |
Dual write | Write to both database and search. Simple but can get out of sync on failures. |
Change Data Capture (CDC) | Capture database changes, stream to search. Reliable, eventually consistent. |
Periodic reindex | Rebuild index on schedule. Simple, but stale between rebuilds. |
Event-driven | Publish events on changes, consumer updates search. Decoupled, scalable. |
Recommended: CDC (using Debezium) or event-driven for production systems.
Design a Product Search System
You're building product search for an e-commerce site with 10 million products. Users search by name, filter by category and price, and sort by relevance or price. How would you design it?
See recommended approach
Architecture:
User search → API Server → Elasticsearch
↘ PostgreSQL (product details)
Product updates → Kafka → Search Indexer → Elasticsearch
Index design:
{
"mappings": {
"properties": {
"name": { "type": "text", "analyzer": "standard" },
"description": { "type": "text" },
"category": { "type": "keyword" },
"price": { "type": "float" },
"rating": { "type": "float" },
"in_stock": { "type": "boolean" },
"brand": { "type": "keyword" },
"tags": { "type": "keyword" }
}
}
}
Query strategy:
- Full-text search on
nameanddescription - Filter by
category,pricerange,in_stock - Boost by
ratingfor relevance - Return IDs, fetch full details from PostgreSQL
Sync strategy:
- Product changes published to Kafka
- Search indexer consumes events, updates Elasticsearch
- Eventually consistent (few seconds delay acceptable)
Scaling:
- 10M products ≈ 10-50 GB index
- Single Elasticsearch cluster sufficient
- Add replicas for read scaling
Search Features
Full-Text Search
{
"query": {
"match": {
"title": "wireless headphones"
}
}
}
Filters (Exact Match)
{
"query": {
"bool": {
"must": { "match": { "title": "headphones" } },
"filter": [
{ "term": { "category": "electronics" } },
{ "range": { "price": { "lte": 200 } } }
]
}
}
}
Fuzzy Search (Typo Tolerance)
{
"query": {
"match": {
"title": {
"query": "headhpones",
"fuzziness": "AUTO"
}
}
}
}
Aggregations (Facets)
{
"aggs": {
"categories": {
"terms": { "field": "category" }
},
"price_ranges": {
"range": {
"field": "price",
"ranges": [
{ "to": 50 },
{ "from": 50, "to": 100 },
{ "from": 100 }
]
}
}
}
}
Returns: counts per category, counts per price range.
Autocomplete/Suggestions
{
"suggest": {
"product-suggest": {
"prefix": "wire",
"completion": {
"field": "suggest"
}
}
}
}
Relevance and Ranking
TF-IDF and BM25
Default relevance scoring based on:
- Term Frequency (TF): How often term appears in document
- Inverse Document Frequency (IDF): How rare is the term
- Field length: Shorter fields with matches score higher
Custom Scoring
Boost results based on business logic:
{
"query": {
"function_score": {
"query": { "match": { "title": "headphones" } },
"functions": [
{
"field_value_factor": {
"field": "rating",
"factor": 1.2
}
},
{
"filter": { "term": { "is_promoted": true } },
"weight": 2
}
]
}
}
}
Level-Based Expectations
| Search System Knowledge by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Know when to use search vs database. Understand basic full-text search concepts. Can explain inverted indexes at high level. |
Senior (L5) | Design index schema for use case. Implement sync strategy (CDC, events). Configure relevance scoring. Handle scaling. |
Staff+ (L6+) | Design search infrastructure at scale. Optimize query performance. Handle multi-tenant search. Consider search quality metrics. |
Engineering Manager | Evaluate build vs buy (Elasticsearch vs Algolia). Plan team expertise needs. Understand search relevance impact on business. |
Common Pitfalls
1. Using Search as Primary Database
Search systems are eventually consistent and not designed for transactional workloads.
Do: Primary database for writes, search index for reads Don't: Use Elasticsearch as your only data store
2. Over-Indexing
Don't index fields you won't search or filter on.
Do: Index searchable fields, store the rest Don't: Index entire documents "just in case"
3. Ignoring Mapping
Schema matters in search. Define mappings explicitly.
{
"mappings": {
"properties": {
"category": { "type": "keyword" }, // Explicit: exact match
"title": { "type": "text" } // Explicit: full-text search
}
}
}
Don't: Rely on dynamic mapping in production
4. Not Measuring Search Quality
Search is never "done." Measure and iterate.
- Click-through rate on search results
- Zero-result searches
- Search refinement rate
- Conversion from search
Search relevance is both art and science. Start simple, measure results, and iterate. Perfect relevance on day one isn't realistic.
Common Interview Scenarios
"How would you implement search for a large catalog?"
"I'd use Elasticsearch as a dedicated search layer. Products would be indexed with searchable fields—name, description, category, tags. The primary PostgreSQL database remains the source of truth. Updates flow through Kafka to a search indexer that updates Elasticsearch. Search queries hit Elasticsearch for fast, relevant results. For the 10 million products scale, a single Elasticsearch cluster with replicas would handle it."
"How do you keep the search index in sync?"
"I'd use change data capture with Debezium. It captures all changes from the PostgreSQL WAL and publishes to Kafka. A search indexer service consumes these events and updates Elasticsearch. This is eventually consistent—a few seconds delay—but reliable. If we need faster sync for critical updates, we could also do a synchronous index update for those specific cases."
"How do you handle typos in search?"
"Elasticsearch has built-in fuzzy matching. I'd configure the search query with fuzziness set to AUTO, which allows for 1-2 character edits based on word length. For autocomplete, I'd use edge n-grams to match partial inputs. Additionally, I could maintain a custom dictionary of common misspellings in our domain."
What's Next
You've now covered the core building blocks for system design. Practice combining them to solve real interview problems in the Questions section.