Blob (Binary Large Object) storage is designed for storing unstructured data like images, videos, documents, and backups. It's a fundamental building block for any system handling media or large files.
Why Blob Storage Matters
Storing large files in your database is problematic:
- Databases aren't optimized for large binary data
- Expensive storage (database storage costs more)
- Backup and replication complexity
- Query performance degradation
Blob storage solves these problems with purpose-built systems optimized for large files.
Rule of thumb: If it's larger than a few KB and doesn't need to be queried, it probably belongs in blob storage.
Blob Storage Technologies
| Popular Blob Storage Solutions | |
|---|---|
| Name | Description |
Amazon S3 | Industry standard. Highly durable (11 9s), scalable, rich ecosystem. The default choice for most systems. |
Google Cloud Storage | GCP equivalent. Similar features to S3. Good integration with GCP services. |
Azure Blob Storage | Microsoft's offering. Strong enterprise features. Best for Azure-native applications. |
MinIO | S3-compatible open source. Good for on-premise or self-hosted needs. |
Cloudflare R2 | S3-compatible with no egress fees. Good for read-heavy workloads. |
Core Concepts
Objects, Buckets, and Keys
Bucket: my-app-images
└── Key: users/123/profile.jpg
└── Key: users/123/cover.jpg
└── Key: products/456/thumbnail.png
- Object: The file plus metadata
- Bucket: Container for objects (like a top-level folder)
- Key: Unique identifier within a bucket (like a file path)
Storage Classes
Different tiers for different access patterns:
| S3 Storage Classes | |
|---|---|
| Name | Description |
Standard | Frequently accessed data. Highest cost, lowest latency. Use for: active user uploads, current content. |
Infrequent Access (IA) | Less frequently accessed. Lower storage cost, retrieval fee. Use for: older content, backups. |
Glacier | Archive storage. Very low cost, minutes-to-hours retrieval. Use for: compliance archives, old backups. |
Intelligent-Tiering | Automatic tiering based on access patterns. Use for: unpredictable access patterns. |
Durability vs Availability
- Durability: Will my data survive? (S3: 99.999999999% - 11 nines)
- Availability: Can I access it right now? (S3 Standard: 99.99%)
Blob storage is extremely durable—you're more likely to have the building burn down than lose data in S3.
Common Patterns
Direct Upload
Client uploads directly to blob storage, bypassing your servers.
1. Client requests upload URL from your API
2. Your API generates pre-signed URL
3. Client uploads directly to S3
4. Client notifies your API of completion
Benefits:
- Reduces load on your servers
- Faster uploads (direct to storage)
- Lower bandwidth costs
# Generate pre-signed upload URL
import boto3
s3 = boto3.client('s3')
url = s3.generate_presigned_url(
'put_object',
Params={'Bucket': 'my-bucket', 'Key': 'uploads/file.jpg'},
ExpiresIn=3600 # URL valid for 1 hour
)
CDN Integration
Serve blob storage content through a CDN.
User → CDN (edge cache) → S3 (origin)
↓
Cached at edge for subsequent requests
Benefits:
- Lower latency globally
- Reduced S3 request costs
- Protection from traffic spikes
Processing Pipeline
Process uploads asynchronously.
Upload → S3 → Event notification → Lambda/Worker
↓
Process (resize, transcode)
↓
Store result → S3
Use case: Image thumbnails, video transcoding, document parsing
Design an Image Upload System
Users upload profile pictures that need to be displayed in multiple sizes (thumbnail, medium, large). How would you design this?
See recommended approach
Architecture:
-
Upload flow:
- Client requests pre-signed URL from API
- Client uploads original image directly to S3
- S3 triggers event notification
-
Processing flow:
- Event triggers image processing worker (Lambda or container)
- Worker downloads original, generates sizes (100px, 400px, 1200px)
- Worker uploads processed images to S3
- Worker updates database with URLs
-
Serving flow:
- Images served through CDN
- URL pattern:
cdn.example.com/users/123/profile-{size}.jpg - Database stores base path, client constructs size-specific URLs
Key decisions:
- Pre-signed URLs: Reduce server load, improve upload speed
- Async processing: User doesn't wait for resizing
- CDN: Fast global delivery, cache at edge
- Naming convention: Predictable URLs, easy to construct
Failure handling:
- If processing fails, retry via dead-letter queue
- Serve original as fallback while processing
Security Considerations
Access Control
| Blob Storage Access Control | |
|---|---|
| Name | Description |
Bucket policies | JSON policies defining who can access what. Applied at bucket level. |
IAM policies | Permissions for AWS users/roles. More flexible, centrally managed. |
Pre-signed URLs | Temporary access to specific objects. Good for controlled sharing. |
Access points | Named network endpoints with distinct permissions. Good for multi-tenant access. |
Pre-Signed URLs
Grant temporary access without exposing credentials.
# Download URL (expires in 1 hour)
url = s3.generate_presigned_url(
'get_object',
Params={'Bucket': 'my-bucket', 'Key': 'private/doc.pdf'},
ExpiresIn=3600
)
# URL: https://bucket.s3.amazonaws.com/private/doc.pdf?X-Amz-Signature=...
Use cases:
- Private content (paid downloads, user documents)
- Secure uploads without exposing credentials
- Time-limited sharing
Encryption
| Encryption Options | |
|---|---|
| Name | Description |
Server-side (SSE-S3) | S3 manages keys. Simplest option, enabled by default. |
Server-side (SSE-KMS) | AWS KMS manages keys. Better audit trail, key rotation. |
Server-side (SSE-C) | Customer provides keys. Full control, more responsibility. |
Client-side | Encrypt before upload. S3 never sees plaintext. Maximum security. |
Performance Optimization
Key Naming for Performance
S3 partitions by key prefix. Previously, sequential keys (timestamps) caused hot spots.
# Old problem (pre-2018, S3 auto-handles now):
uploads/2024-01-01-00001.jpg # All hit same partition
uploads/2024-01-01-00002.jpg
# Still good practice for extreme scale:
uploads/a1b2/2024-01-01-00001.jpg # Hash prefix distributes
uploads/c3d4/2024-01-01-00002.jpg
Multipart Uploads
For large files (>100MB), upload in parts.
1. Initiate multipart upload
2. Upload parts in parallel (e.g., 10MB chunks)
3. Complete multipart upload
Benefits:
- Parallel upload = faster
- Resume failed uploads
- Required for files > 5GB
Transfer Acceleration
For global users uploading to a specific region.
User in Tokyo → Edge location → AWS backbone → S3 in US
(faster than)
User in Tokyo → Public internet → S3 in US
Level-Based Expectations
| Blob Storage Knowledge by Level | |
|---|---|
| Name | Description |
Mid-Level (L4) | Know to use blob storage for media/files. Understand basic upload/download patterns. Can mention S3 as the go-to solution. |
Senior (L5) | Design pre-signed URL flows. Integrate with CDN. Discuss storage classes and lifecycle policies. Handle async processing. |
Staff+ (L6+) | Optimize for extreme scale (multipart, key naming). Design multi-region replication. Discuss cost optimization strategies. Handle compliance requirements. |
Engineering Manager | Evaluate storage costs at scale. Plan data lifecycle and retention policies. Understand compliance implications (data residency, encryption). |
Common Interview Scenarios
"How do you handle user uploads?"
"For user uploads, I'd use pre-signed URLs to allow direct uploads to S3, bypassing our servers. When the upload completes, an S3 event triggers a Lambda function for any processing needed—like generating thumbnails or scanning for malware. The processed files are served through CloudFront CDN. The database stores references to the S3 keys, not the files themselves."
"How do you handle very large files?"
"For large files, I'd use multipart upload. The file is split into chunks (say, 10MB each), uploaded in parallel, then assembled by S3. This allows resuming failed uploads and is faster due to parallelization. For files over 5GB, multipart is required."
"How do you control access to private files?"
"I'd use pre-signed URLs with short expiration times. When a user requests access to their document, my API verifies their authorization, then generates a pre-signed URL valid for 5-15 minutes. This way, the user gets temporary direct access without my server proxying the file."
What's Next
Blob storage handles file storage. Next, we'll look at CDNs, which cache and serve content at the edge for faster global delivery.