NEW: ML Mock & Coaching now available

Building Blocks

Blob Storage

Understanding object storage systems like S3, when to use them, and how to integrate them in system design.

7 min read

Blob (Binary Large Object) storage is designed for storing unstructured data like images, videos, documents, and backups. It's a fundamental building block for any system handling media or large files.

Why Blob Storage Matters

Storing large files in your database is problematic:

  • Databases aren't optimized for large binary data
  • Expensive storage (database storage costs more)
  • Backup and replication complexity
  • Query performance degradation

Blob storage solves these problems with purpose-built systems optimized for large files.

Info

Rule of thumb: If it's larger than a few KB and doesn't need to be queried, it probably belongs in blob storage.

Blob Storage Technologies

Popular Blob Storage Solutions
NameDescription
Amazon S3Industry standard. Highly durable (11 9s), scalable, rich ecosystem. The default choice for most systems.
Google Cloud StorageGCP equivalent. Similar features to S3. Good integration with GCP services.
Azure Blob StorageMicrosoft's offering. Strong enterprise features. Best for Azure-native applications.
MinIOS3-compatible open source. Good for on-premise or self-hosted needs.
Cloudflare R2S3-compatible with no egress fees. Good for read-heavy workloads.

Core Concepts

Objects, Buckets, and Keys

Bucket: my-app-images
  └── Key: users/123/profile.jpg
  └── Key: users/123/cover.jpg
  └── Key: products/456/thumbnail.png
  • Object: The file plus metadata
  • Bucket: Container for objects (like a top-level folder)
  • Key: Unique identifier within a bucket (like a file path)

Storage Classes

Different tiers for different access patterns:

S3 Storage Classes
NameDescription
StandardFrequently accessed data. Highest cost, lowest latency. Use for: active user uploads, current content.
Infrequent Access (IA)Less frequently accessed. Lower storage cost, retrieval fee. Use for: older content, backups.
GlacierArchive storage. Very low cost, minutes-to-hours retrieval. Use for: compliance archives, old backups.
Intelligent-TieringAutomatic tiering based on access patterns. Use for: unpredictable access patterns.

Durability vs Availability

  • Durability: Will my data survive? (S3: 99.999999999% - 11 nines)
  • Availability: Can I access it right now? (S3 Standard: 99.99%)

Blob storage is extremely durable—you're more likely to have the building burn down than lose data in S3.

Common Patterns

Direct Upload

Client uploads directly to blob storage, bypassing your servers.

1. Client requests upload URL from your API
2. Your API generates pre-signed URL
3. Client uploads directly to S3
4. Client notifies your API of completion

Benefits:

  • Reduces load on your servers
  • Faster uploads (direct to storage)
  • Lower bandwidth costs
# Generate pre-signed upload URL
import boto3

s3 = boto3.client('s3')
url = s3.generate_presigned_url(
    'put_object',
    Params={'Bucket': 'my-bucket', 'Key': 'uploads/file.jpg'},
    ExpiresIn=3600  # URL valid for 1 hour
)

CDN Integration

Serve blob storage content through a CDN.

User → CDN (edge cache) → S3 (origin)
      ↓
      Cached at edge for subsequent requests

Benefits:

  • Lower latency globally
  • Reduced S3 request costs
  • Protection from traffic spikes

Processing Pipeline

Process uploads asynchronously.

Upload → S3 → Event notification → Lambda/Worker
                                      ↓
                                  Process (resize, transcode)
                                      ↓
                                  Store result → S3

Use case: Image thumbnails, video transcoding, document parsing

Challenge

Design an Image Upload System

Users upload profile pictures that need to be displayed in multiple sizes (thumbnail, medium, large). How would you design this?

See recommended approach

Architecture:

  1. Upload flow:

    • Client requests pre-signed URL from API
    • Client uploads original image directly to S3
    • S3 triggers event notification
  2. Processing flow:

    • Event triggers image processing worker (Lambda or container)
    • Worker downloads original, generates sizes (100px, 400px, 1200px)
    • Worker uploads processed images to S3
    • Worker updates database with URLs
  3. Serving flow:

    • Images served through CDN
    • URL pattern: cdn.example.com/users/123/profile-{size}.jpg
    • Database stores base path, client constructs size-specific URLs

Key decisions:

  • Pre-signed URLs: Reduce server load, improve upload speed
  • Async processing: User doesn't wait for resizing
  • CDN: Fast global delivery, cache at edge
  • Naming convention: Predictable URLs, easy to construct

Failure handling:

  • If processing fails, retry via dead-letter queue
  • Serve original as fallback while processing

Security Considerations

Access Control

Blob Storage Access Control
NameDescription
Bucket policiesJSON policies defining who can access what. Applied at bucket level.
IAM policiesPermissions for AWS users/roles. More flexible, centrally managed.
Pre-signed URLsTemporary access to specific objects. Good for controlled sharing.
Access pointsNamed network endpoints with distinct permissions. Good for multi-tenant access.

Pre-Signed URLs

Grant temporary access without exposing credentials.

# Download URL (expires in 1 hour)
url = s3.generate_presigned_url(
    'get_object',
    Params={'Bucket': 'my-bucket', 'Key': 'private/doc.pdf'},
    ExpiresIn=3600
)
# URL: https://bucket.s3.amazonaws.com/private/doc.pdf?X-Amz-Signature=...

Use cases:

  • Private content (paid downloads, user documents)
  • Secure uploads without exposing credentials
  • Time-limited sharing

Encryption

Encryption Options
NameDescription
Server-side (SSE-S3)S3 manages keys. Simplest option, enabled by default.
Server-side (SSE-KMS)AWS KMS manages keys. Better audit trail, key rotation.
Server-side (SSE-C)Customer provides keys. Full control, more responsibility.
Client-sideEncrypt before upload. S3 never sees plaintext. Maximum security.

Performance Optimization

Key Naming for Performance

S3 partitions by key prefix. Previously, sequential keys (timestamps) caused hot spots.

# Old problem (pre-2018, S3 auto-handles now):
uploads/2024-01-01-00001.jpg  # All hit same partition
uploads/2024-01-01-00002.jpg

# Still good practice for extreme scale:
uploads/a1b2/2024-01-01-00001.jpg  # Hash prefix distributes
uploads/c3d4/2024-01-01-00002.jpg

Multipart Uploads

For large files (>100MB), upload in parts.

1. Initiate multipart upload
2. Upload parts in parallel (e.g., 10MB chunks)
3. Complete multipart upload

Benefits:
- Parallel upload = faster
- Resume failed uploads
- Required for files > 5GB

Transfer Acceleration

For global users uploading to a specific region.

User in Tokyo → Edge location → AWS backbone → S3 in US
              (faster than)
User in Tokyo → Public internet → S3 in US

Level-Based Expectations

Blob Storage Knowledge by Level
NameDescription
Mid-Level (L4)Know to use blob storage for media/files. Understand basic upload/download patterns. Can mention S3 as the go-to solution.
Senior (L5)Design pre-signed URL flows. Integrate with CDN. Discuss storage classes and lifecycle policies. Handle async processing.
Staff+ (L6+)Optimize for extreme scale (multipart, key naming). Design multi-region replication. Discuss cost optimization strategies. Handle compliance requirements.
Engineering ManagerEvaluate storage costs at scale. Plan data lifecycle and retention policies. Understand compliance implications (data residency, encryption).

Common Interview Scenarios

"How do you handle user uploads?"

"For user uploads, I'd use pre-signed URLs to allow direct uploads to S3, bypassing our servers. When the upload completes, an S3 event triggers a Lambda function for any processing needed—like generating thumbnails or scanning for malware. The processed files are served through CloudFront CDN. The database stores references to the S3 keys, not the files themselves."

"How do you handle very large files?"

"For large files, I'd use multipart upload. The file is split into chunks (say, 10MB each), uploaded in parallel, then assembled by S3. This allows resuming failed uploads and is faster due to parallelization. For files over 5GB, multipart is required."

"How do you control access to private files?"

"I'd use pre-signed URLs with short expiration times. When a user requests access to their document, my API verifies their authorization, then generates a pre-signed URL valid for 5-15 minutes. This way, the user gets temporary direct access without my server proxying the file."

What's Next

Blob storage handles file storage. Next, we'll look at CDNs, which cache and serve content at the edge for faster global delivery.