AO
Back
Harvey AI Common Problems

Google Drive / File Storage System Design with ACL

System DesignhardLast reported July 2026
By AceOffer · Updated July 2026 · Reported 11× across 50+ reports

Understanding the Problem

Design a Google Drive-like file storage system. Core functional requirements: (1) file listing, (2) file upload and download, (3) file/folder sharing, (4) access control. The interviewer typically starts with a base scope and may add requirements during clarification. One common variant ('Data Room') restricts to PDF files only, emphasizes organization-level ACL for sharing with thousands of users, and specifies ~100k total users with scalability not a primary constraint. Other variants emphasize scaling deep-dive, security, or all four dimensions equally. Multi-device syncing is explicitly out of scope in several instances. The system should support both user-level and organization/group-level permission management.

Functional Requirements

Structured requirements coming soon. For now, see the full problem statement above and the deep-dive prompts below.

Non-Functional Requirements

Latency, throughput, availability, consistency targets — being authored.

The Set Up

Defining the Core Entities

Core entities (Request, Batch, Worker, Cache, etc.) — being authored.

The API

POST /endpoint → describe request shape GET /endpoint → describe response shape (API spec being authored)

High-Level Design

Component diagram + walkthrough mapping each functional requirement to a system flow — being authored.

Potential Deep Dives

These are the directions the interviewer is likely to push you. Each one has multiple valid solutions at different quality tiers.

1)How do you handle very large files? (when: Candidate presents basic upload design)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

2)How do you ensure/verify that each uploaded chunk actually belongs to the user who requested the upload? (when: Candidate mentions chunked upload to S3)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

3)What if a file needs to be shared with an entire organization of thousands of users? (when: Candidate designs user-level ACL)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

4)How do you handle folder-level permissions and inheritance? (when: Candidate presents flat folder/file structure)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

5)How do you handle security for file access? (when: Candidate presents basic design without security discussion)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

6)How do you scale this system? (when: System design at scale)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

7)How do you verify file consistency / detect duplicate content? (when: Candidate mentions file deduplication or integrity)

Bad

Naive approach with serious trade-off — being authored.

Good

Solid baseline with reasonable trade-offs — being authored.

Great

Production-grade approach with explicit trade-off rationale — being authored.

What is Expected at Each Level?

L4 / Mid-level

Cover happy path. Clarify scope. Identify the obvious bottleneck. Pick a reasonable storage and reasonable scaling approach.

L5 / SeniorTarget

All of the above plus: explicit failure handling, durability vs latency trade-offs, choose the right batching/caching strategy, articulate why.

L6 / Staff+

All of the above plus: organizational concerns (rollout, migration, on-call), quantitative analysis, multi-region considerations, what could go wrong with the proposed solution at 10x scale.

Insider Notes

Common mistakes: Not proactively raising chunking for large files — waiting for interviewer to bring it up; Designing only user-level ACL without considering organization/group-level sharing at scale; Spending too long on requirements clarification, leaving insufficient time for deep design; Not addressing security (pre-signed URLs, access token scoping) unless prompted; Getting overwhelmed by scope when interviewer adds too many features during clarification; Treating folder permission inheritance as trivial without discussing override semantics

Interviewer hints: When candidate proactively mentioned chunking, interviewer was visibly satisfied (confirmed as good approach); Interviewer framed follow-up directions as four areas: scaling, access control (+ 2 others not recalled); One interviewer (new/inexperienced) kept confirming support for every feature asked — candidate should have pushed back to scope down

What passers do: Clearly separated metadata store (relational DB) from blob store (S3) early in the design; Proactively introduced chunking for large files before being asked; Addressed both user-level and organization/group-level ACL; Described pre-signed URL flow for both upload and download; Explained chunk ownership verification (pre-signed URL scoped to specific S3 path); Kept design concise and invited follow-up rather than trying to cover everything upfront

Why people fail: Failed to cover ACL design in sufficient depth; Did not address large file handling without explicit prompt; Ran over time due to excessive clarification or over-explaining early sections; Designed monolithic storage without blob/metadata separation; Could not answer how to verify chunk ownership when probed

Edge cases probed: Sharing a file with an organization of thousands of users — O(N) ACL rows vs. single org-level entry; Large file upload: chunk ownership verification via pre-signed S3 URL scoped to specific key; File consistency check / integrity verification across chunks (hash per chunk + full file hash); Folder permission inheritance vs. explicit overrides at child level; Duplicate file names in same directory (e.g., 'A' → 'A(1)' on re-insertion); Scope creep: interviewer keeps adding features during clarification — candidate must time-box

Alternative approaches: Flat ACL per user (instead of group/org-level) (Simple to query but doesn't scale when sharing with thousands of users in an org — requires O(N) rows per share operation. Group/org-level ACL with membership lookup is preferred.); Trie-based in-memory filesystem for metadata (Good for coding-round filesystem simulation; not suitable for production persistence. Useful for the companion coding question but not the system design.); Monolithic storage (no blob/metadata split) (Simpler architecture but does not scale; storing binary data in relational DB is inefficient and costly for large files.)

Harvey AI · System Design · Last reported July 2026
Is this helpful?