NEW: ML Mock & Coaching now available

Questions

Meta Post Privacy

MetaTikTok

Design a scalable privacy system for Meta posts - supporting Friends, Friends of Friends, custom audience controls, and global enforcement across feeds, search, and CDN media delivery.

35 min read

Challenge

Think Beyond the Happy Path

Real-world systems like Meta's post privacy aren't just running short tasks on schedule. They're about billions of access checks per day, exploding friend-of-friend graphs, and cache invalidations that ripple across millions of users.

Before we dive into solutions, pause and ask:

  • How do you keep privacy checks fast at Meta scale?
  • How do you handle a Friends of Friends (FOF) set that grows to tens of millions?
  • What happens when a single friendship change forces cache invalidation across the globe?
  • How do you design revocation that instantly invalidates access across feeds, search, and CDN media delivery?

You don't need to answer all of these right away. A Senior/Staff+ Engineer doesn't stop at basic functionality — they anticipate edge cases, design for resilience, and push for production-grade reliability.

But great systems start with great questions. What would you ask next?

Problem Statement

FB post privacy is about choosing who gets to see your content—friends, everyone, or only yourself—on a post-by-post basis, with options to fine-tune visibility.

Functional Requirements

  • FR1: Standard Privacy Options

    Authors can choose from basic visibility modes like Only Me, Friends, or Public.

  • FR2: Extended & Custom Audience Controls

    More advanced rules must be supported, such as Friends of Friends (FOF) or Friends except…, including inclusion and exclusion lists.

  • FR3: Global Enforcement

    Rules must hold consistently everywhere: feeds, profiles, search, notifications, sharing, and even media served from CDNs.

Non-Functional Requirements

  • NFR1: Low Latency — Checks should complete in single-digit milliseconds at P99.
  • NFR2: High Availability — 99.99% uptime, with a principle of "fail closed" (deny if uncertain).
  • NFR3: Scalability — The policy decision, graph lookups, and caches must scale horizontally in Billion user & connection level.
  • NFR4: Correctness — Eventual consistency is acceptable, but revocations must converge fast across the system.
Requirements Summary
NameDescription
FR1Standard Privacy Options | Authors can choose from basic visibility modes like Only Me, Friends, or Public.
FR2Extended & Custom Audience Controls | More advanced rules such as Friends of Friends (FOF), Friends except…, with inclusion and exclusion lists.
FR3Global Enforcement | Rules must hold consistently everywhere: feeds, profiles, search, notifications, sharing, and CDN media.
NFR1Low Latency | Single-digit milliseconds at P99 for privacy checks.
NFR2High Availability | 99.99% uptime with fail-closed behavior (deny if uncertain).
NFR3Scalability | Policy decision, graph lookups, and caches scale horizontally at billion-user level.
NFR4Correctness | Eventual consistency acceptable, but revocations must converge fast across the system.

Core Concept: Audience as Set Algebra

Think of post visibility like a filter: we start with everyone, then add rules to decide who's in and who's out.

Step 1: Base Audience

When you create a post, you pick a base rule:

  • Only Me → just the author.
  • Friends → author's friends.
  • Friends of Friends (FOF) → author's friends, plus their friends.
  • Public → anyone.

Step 2: Includes & Excludes

Then, the author can fine-tune:

  • Include lists or specific people → always add them in.
  • Exclude lists or specific people → always keep them out.

Step 3: Final Check

To decide if a viewer V can see a post by author A:

  1. Start with the base audience.
  2. Add people from Includes.
  3. Remove people from Excludes and any blocked users.
  4. If V is still in → they can see it. Otherwise → hidden.
Info

Why this matters?

By thinking in terms of "start big, then filter", complex cases like custom lists or friends of friends become much easier. Under the hood, the system just needs fast ways to check:

  • "Is this person in the allowed set?"
  • "Are they excluded or blocked?"

That's how Meta can answer billions of these checks every day at scale.

API & Schema Sketch

1) Create/Update Audience Rules: Normalize privacy once; reuse across posts.

  • Create: POST /v1/audiencespecs

    { "base": "FRIENDS", "includes": {"list_ids":["L1"],"user_ids":["u42"]}, "excludes": {"list_ids":[], "user_ids":["u77"]} }
    

    → 201 { "audience_id": "A1", "version": 1 }

  • Update (versioned): PATCH /v1/audiencespecs/{audience_id}

    → 200 { "audience_id":"A1","version":2 }

    Key validations needed: caller owns lists; no user/list in both include & exclude.

2) Create Post (attach audience): Source of truth tying content to privacy.

  • Create: POST /v1/posts

    { "author_id":"u1", "content":{"text":"hello"}, "audience_id":"A1" }
    

    → 201 { "post_id":"P1","audience_id":"A1","created_at":"..." }

  • (Optional) Change audience: PATCH /v1/posts/{post_id} with { "audience_id":"A2" }

    Key validations needed: author = caller; audience exists; audit the change.

3) Batch Visibility Check: Fast feed/search filtering is impossible without this.

  • Check: POST /v1/visibility:batch

    { "viewer_id":"u42", "post_ids":["P1","P2","P3"] }
    

    →

    { "results": [
      {"post_id":"P1","visible":true,"reason":"IN_BASE_FRIENDS"},
      {"post_id":"P2","visible":false,"reason":"EXCLUDED_BY_LIST"},
      {"post_id":"P3","visible":true,"reason":"IN_INCLUDES"}
    ]}
    

Notes: Server expands graph/lists once per viewer; returns reason codes for debugging.

Core Entity

At the heart of a privacy-aware social system are a handful of core entities. Each table plays a very specific role in anchoring identity, maintaining the social graph, encoding privacy rules, and binding them to actual content. The key principle here is separation of concerns: keep identity, relationships, privacy rules, and posts distinct, but link them together so that visibility can be evaluated consistently and retroactively.

Users

Users
NameDescription
user_id (PK)Stable identity; never reused. Drives ownership and audit trails.
nameDisplay only; not used for joins.
created_atFor audit/forensics and abuse heuristics.

Friendships

Design intent: Compact, deterministic undirected graph for FRIENDS/FOF checks.

Friendships
NameDescription
u (PK part)Smaller user id; enforce u < v to store each undirected edge once.
v (PK part)Larger user id.
created_atUseful for abuse detection and FOF TTL heuristics.

Audience_specs

Design intent: Store rules, not expanded members so changes propagate system‑wide without rewrites.

Audience_specs
NameDescription
audience_id (PK)Stable id for reuse across posts.
owner_id (FK → users.user_id)Only owner may reference/modify. Validated at write.
versionMonotonic; enables retroactive policy edits vs historical snapshots.
baseENUM: ONLY_ME, FRIENDS, FOF, PUBLIC.
includesLogical references (list ids, user ids). Keep compact; no expansion at write.
excludesSame shape as includes; excludes win on conflict.
created_atFor audit and cache versioning/epochs.

Posts

Design intent: Bind content to a policy reference (not a member list) so the privacy check can re‑evaluate on every read.

Posts
NameDescription
post_id (PK)Content identity.
author_id (FK → users.user_id)Used for blocks and same‑owner shortcuts.
audience_id (FK → audience_specs.audience_id)Rule reference.
audience_versionNULL = FOLLOW latest; set = PINNED snapshot.
content_jsonMetadata pointer; bytes live in object storage guarded by signed URLs.
created_atOrdering, feeds, retention.
updated_atFor edits and invalidation signals.

High-Level Architecture

FR1: Standard Privacy Options

Start with "Friends" and "Only Me" posts to establish the core privacy loop. By avoiding FOF and custom lists at first, you skip graph blow-ups and list expansion, and focus on fundamentals you'll reuse for all modes: record rule → attach to post → batch check → deny-first → allow/deny → audit → (later) propagate changes.

1.1 Define Privacy via audience_spec_id

Privacy Definition Flow

Click to expand

At this point, user chooses a privacy rule in the UI before submitting a post, with following privacy options include:

  • Only Me
  • Friends
  • Friends of Friends (FOF)
  • Public
  • Custom (e.g., "Friends except John", "Visible to List A + User B")

The frontend converts this into a structured AudienceSpec, for example:

{
  "base": "FRIENDS",
  "includes": { "user_ids": ["u777"], "list_ids": ["L123"] },
  "excludes": { "user_ids": ["u999"] }
}
Rule-Based Privacy (AudienceSpec) VS Flat User Lists

Storing rule-based privacy configs (like audience_spec_id) instead of expanding to exact user IDs at write time is one of the most important architectural decisions in systems like Meta. Here's the why, explained step-by-step.

1. Rules Stay Small & Reusable — Users Lists Can Explode

A privacy rule like below: is typically a few kilobytes — even for large authors.

{ base: "FRIENDS", includes: [...], excludes: [...] }

If you expanded this rule into actual user IDs (including friends, friends-of-friends, list members), you might store thousands or millions of user IDs per post.

  • A celebrity with 5,000 friends might generate a FOAF set with 10M+ users.
  • That's huge write amplification, especially on mobile.

Storing the rule (not the result) keeps writes cheap and avoids duplicating logic per post.

2. Support Retroactive Privacy Changes

When you store the rule, you can re-evaluate it at read time.

That means:

  • If the author blocks someone, they're excluded immediately.
  • If they remove someone from a custom list, that user loses access.
  • If the author changes the audience base from FRIENDS to ONLY_ME, it revokes access to everyone else, even retroactively.

We will discuss about this in the Deep Dive 2

3. Policy Uniformity Across Surfaces

Using rule-based configs ensures the same privacy logic applies:

  • On the feed
  • On the profile page
  • On permalinks
  • On search results
  • On the CDN when fetching media

One audience_spec_id + version → evaluated the same way everywhere

We will discuss about this in the FR3

1.2 Step-by-step data flow — Create Post (write path)

Let's walk through what actually happens when a user presses "Post" in a system like Meta. The key idea is that we don't just dump content into a database. We carefully preserve intent, validate privacy rules, ensure idempotency, and guarantee that media never leaks without proper authorization.

Step 1 — Client → API: Capture Intent

The client sends a POST /posts { content, audience_spec_id } request to the API Gateway. At this stage, the system only captures intent. Nothing is visible yet.

The API Gateway does some critical checks before we even think about writing anything:

  • Authenticate the request using a JWT or session token.
  • Enforce quotas and limits to prevent abuse (post size, MIME types, number of attachments).

Step 2 — Resolve & Validate Privacy Rule

At this stage, the API calls into the Privacy Service to validate the requested audience_spec_id. This step is critical because without it, a user could attempt to reference someone else's list or craft an invalid rule that bypasses privacy enforcement. Yes, it adds an extra hop to the write path, but the payoff is that we centralize policy enforcement in one canonical place. Failing fast here is far better than persisting a post that later turns out to violate rules. The service checks two main things:

  • Does the author actually own the audience spec.
  • Are the include/exclude lists valid and within system limits.

Once those checks pass, the Privacy Service responds with a normalized, authoritative spec object, for example:

{
  "owner_id": "u123",
  "current_version": 3,
  "base": "FRIENDS",
  "limits_ok": true,
  "ownership_ok": true
}

Step 3 — Store Post into DB

Once the privacy spec has been validated by the Privacy Service, control shifts to the Content Service, which is responsible for persisting the post itself.

Post Storage Flow

Click to expand

At this stage, the system needs to record not only the post content but also a durable link back to the privacy rule that governs it. This is done by writing the post with (author_id, audience_spec_id) .

Post Audience Spec Versioning?

When a post is persisted, the system needs to decide how tightly it should bind to its privacy rule. We support two versioning modes:

  • FOLLOW mode → The post stores audience_ver = NULL. This means the post always resolves to the latest version of the audience spec at read time. If the author later edits the privacy rule (e.g., adds someone to "friends except"), those changes automatically apply to all FOLLOW-mode posts.
  • PINNED mode → The post stores audience_ver = resolved_version. In this case, the post is effectively "frozen" to the exact snapshot of the rule at the moment of creation. Even if the author edits the rule later, this post continues to respect the older version.

Next, media attachments are handled. Images or videos that were uploaded earlier follow a two-phase protocol:

  1. The client requests an upload_id.
  2. The client streams bytes to storage with that ID.
  3. The client finalizes the upload.

By the time the post is created, the Content Service simply binds those uploads to the new post_id. Crucially, all media objects remain non-public during this process: there are no open ACLs. They can only be fetched via signed, short-lived URLs that the Privacy Gateway issues after verifying the viewer's access rights. This ensures that media is never accidentally exposed outside the privacy boundary of the post itself. Details will be discussed in FR3.

1.3: Reading a Post in a Thread — Privacy-Aware Data Flow

Once a post is created and persisted with its (author_id, audience_spec_id, audience_ver) binding, the next challenge is enforcing privacy when someone tries to read it. The read path is where policies actually get exercised: if you fail here, sensitive content leaks.

Step-by-step data flow — Read Post

When a viewer requests to see a post inside a thread — whether they arrived from the feed, a profile, a permalink, or a notification — the system must do more than just pull bytes from storage. It has to carefully enforce privacy at every step of the way. Let's walk through the entire read path.

Step 1 — API Gateway: Capture the request

The journey begins when the client sends a request:

GET /posts/{post_id}?viewer_id=V

The API Gateway sits at the edge, acting as both gatekeeper and traffic cop. It authenticates the caller via JWT, applies rate limits to guard against abuse, and annotates the request with the viewer_id. At this stage we are simply saying, "Viewer V is asking to read Post P." No content is returned yet — all we've done is establish identity and enforce basic guardrails.

Step 2 — Content Service: Retrieve metadata

Next, the API hands the request to the Content Service, which retrieves lightweight metadata for the post. This includes the author_id, the audience_spec_id.

This metadata acts like a pointer — it doesn't yet answer "can Viewer V see this post?" but it tells us where the rules live that will decide the answer.

Step 3 — API → Privacy Service: visibility check

Now comes the heart of the read path: the Privacy Service. The API forwards the tuple (viewer_id, author_id, audience_spec_id, audience_ver), and the Privacy Service orchestrates a series of checks behind the scenes:

  • Fri: Is the viewer a friend of the author? A friend-of-friend? Blocked?
  • Privacy Service: Expand the includes and excludes from the audience_spec.
  • Decision Logic: Combine these inputs into a final allow/deny answer.

Because this logic runs for billions of reads per day, caching is critical. Friend sets are cached per user; audience expansions are cached per (spec, version). For extremely popular authors — think celebrities with millions of followers — the system may need special handling like pre-materialized expansions or approximate sets (e.g., Bloom filters) to keep latency predictable.

Step 4 — Privacy Service: Return the decision

The Privacy Service responds with the verdict. If the viewer is allowed, the API instructs the Content Service to fetch and return the full post body. If the viewer is denied, the system responds with a 403 or a sanitized placeholder such as "This post is not available." At this stage, access is decided for the metadata, but there's still one final piece to guard: the media.

Step 5 — Media Privacy Enforcement

Media objects (images, videos) must never be exposed directly. Even if a viewer has passed the metadata check, the system enforces one more layer:

  • The Privacy Service mints a short-lived signed URL tied to {viewer_id, post_id, exp, audience_epoch}.
  • These URLs are the only way the CDN will serve media bytes.
  • If the author later revokes access — say by blocking the viewer or tightening the post's privacy — simply incrementing audience_epoch immediately invalidates all previously issued tokens at the CDN edge.

This design avoids slow per-object ACL updates and ensures that media distribution honors the same privacy rules as metadata.

Sign in to continue reading

"Meta Post Privacy" requires a free account to access.

Sign in to continue