System Design Interview: How Would You Let Users Upload Huge Files Even If the Internet Disconnects…
Uploading a file sounds simple.
System Design Interview: How Would You Let Users Upload Huge Files Even If the Internet Disconnects Midway?
Uploading a file sounds simple.
Select file.
Click Upload.
Wait for completion.
But things become much more interesting when the file is:
5 GB
or
50 GB
or even
500 GB
Full story for non-members | E-Books on Java/Microservices/Springboot | Whatsapp Group | Instagram-@arvind.codefarm

Now imagine a user is uploading a 5 GB video.
After:
4.8 GB
has already been uploaded, their internet connection drops.
Do we really want to tell them:
“Sorry. Please start again from the beginning.”
Most users would simply close the browser and never return.
Platforms like Google Drive, Dropbox, YouTube, OneDrive, and AWS S3 don’t work that way.
They allow uploads to continue from exactly where they stopped.
Let’s explore how.
The Question
Aadvik: Imagine we’re building something like Google Drive.
A user uploads a 5 GB video.
After 4.8 GB has been uploaded, the network disconnects.
How would you ensure the upload can resume instead of starting from scratch?
Neha: Before discussing solutions, I’d first challenge the assumption.
The problem isn’t the network failure.
The problem is that we’re treating the entire file as a single upload operation.
Aadvik: What do you mean?
Neha: Most naive implementations do something like:
POST /upload
File = 5GB
The server expects the entire file in one request.
If the connection breaks:
Request Failed
Everything is lost.
The upload must restart.
Aadvik: So what’s the alternative?
Neha: Don’t upload the file.
Upload pieces of the file.
Breaking the File into Chunks
Aadvik: Explain.
Neha: Instead of sending:
5 GB File
as a single request,
we split it into smaller chunks.
For example:
Chunk 1 = 8 MB
Chunk 2 = 8 MB
Chunk 3 = 8 MB
...
Chunk N = 8 MB

Now every chunk becomes an independent upload.
Why Is This Better?
Aadvik: Why does chunking help?
Neha: Because failures become smaller.
Imagine:
Chunk 1 Uploaded
Chunk 2 Uploaded
Chunk 3 Uploaded
Chunk 4 Failed
Only Chunk 4 must be retried.
Not the entire 5 GB file.
After reconnecting:
Resume From Chunk 4
instead of:
Restart Entire Upload
Upload Session
Aadvik: How does the server know these chunks belong to the same file?
Neha: We introduce an upload session.
The client first creates an upload session.

Example:
UPLOAD-12345
Every chunk now carries:
UploadSessionId
ChunkNumber
For example:
{
"uploadSessionId":"UPLOAD-12345",
"chunkNumber":42
}
The server can now track progress.
Where Is Progress Stored?
Aadvik: What exactly does the server store?
Neha: Typically metadata.
Something like:
{
"sessionId":"UPLOAD-12345",
"fileName":"video.mp4",
"totalChunks":640,
"uploadedChunks":[1,2,3,4,5]
}
This allows the system to know:
- Which chunks exist
- Which chunks are missing
- Whether upload is complete
The Interview Trap
Aadvik: Let’s say the user refreshes the browser.
What happens now?
Neha: Good question.
If progress only exists in browser memory:
Upload Lost
We have a problem.
Instead, the Upload Session ID must be persisted.
Common options:
Local Storage
Database
Redis
Upload Metadata Store
Now even after:
- Browser refresh
- App restart
- Laptop reboot
the client can continue using the same upload session.
Resume Logic
Aadvik: Walk me through the resume flow.
Neha: Suppose:
640 Chunks Total
The network disconnects after:
Chunk 512
When the client reconnects:

Only missing chunks are uploaded.
Everything else is skipped.
What About Duplicate Chunks?
Aadvik: What if Chunk 513 was uploaded successfully but the acknowledgment got lost?
The client retries.
Neha: That’s a classic distributed systems problem.
The client doesn’t know whether:
Chunk Uploaded
or
Chunk Failed
The server must make chunk uploads idempotent.
Aadvik: How?
Neha: Every chunk should have:
UploadSessionId
ChunkNumber
as a unique identifier.
UPLOAD-12345
Chunk-513
If the same chunk arrives twice:
Ignore Duplicate
or
Return Success Again
without storing it twice.
Exactly the same principle we discussed for payment idempotency.
Storage Layer
Aadvik: Where are these chunks stored?
Neha: Usually object storage.
Examples:
- S3
- GCS
- Azure Blob Storage

The Upload Service manages metadata.
The actual bytes live in object storage.
The Next Scaling Problem
Aadvik: Let’s say 100,000 users are uploading videos simultaneously.
Would all uploads pass through application servers?
Neha: Ideally no.
Application servers would become a bottleneck.
Aadvik: What’s the alternative?
Neha: Direct uploads.
The application generates a temporary upload URL.
The browser uploads directly to object storage.

This is how many cloud-native systems work today.
The application handles metadata.
Storage handles the file transfer.
Chunk Size Selection
Aadvik: How do we choose chunk size?
Neha: That’s always a tradeoff.
Small chunks:
More Requests
More Metadata
Better Recovery
Large chunks:
Fewer Requests
Less Overhead
More Rework On Failure
Common sizes:
5 MB
8 MB
16 MB
64 MB
depending on workload.
File Assembly
Aadvik: Once all chunks arrive, what happens?
Neha: The system assembles them.

Many cloud storage providers perform this merge operation automatically.
Data Integrity
Aadvik: What if a chunk gets corrupted during transmission?
Neha: We verify integrity.
Every chunk typically contains:
Checksum
Hash
ETag
Example:
SHA-256
The server validates the chunk before accepting it.
If validation fails:
Retry Chunk
instead of accepting corrupted data.
Multi-Region Uploads
Aadvik: Let’s make this global.
Users upload from:
- India
- Europe
- US
Any concerns?
Neha: Latency.
Uploading a 10 GB file across continents is painful.
Most platforms use:

Uploads happen close to users.
Replication occurs later.
The Real Production Architecture
Aadvik: If you were designing this today, what would your architecture look like?
Neha:

Every component solves a specific problem:
- Session tracking enables resumability.
- Chunking isolates failures.
- Object storage scales file transfer.
- Metadata enables recovery.
- Checksums ensure integrity.
- Assembly creates the final file.
Lets Conclude
Aadvik: Summarize your solution.
Neha:
- Split large files into chunks.
- Create an Upload Session ID.
- Track uploaded chunks.
- Resume only missing chunks after failure.
- Make chunk uploads idempotent.
- Store files in object storage.
- Use checksums for validation.
- Prefer direct uploads at scale.
- Design assuming networks will fail.
The challenge isn’t uploading large files.
The challenge is ensuring users never lose progress when networks inevitably fail.
That’s why modern platforms rely on chunked uploads, resumable uploads, upload sessions, and object storage to deliver a seamless upload experience.
Liked this deep dive story? If Yes Please 👏 Clap(50) | 📤 Share | 🔔 Follow
Below is a collection of all related stories in one place
https://codefarm0.medium.com/list/15-system-design-interview-scenarios-23f298ce71ad
메타데이터
- post_id
- cab5a3b0abae
- slug
- system-design-interview-how-would-you-let-users-upload-huge-files-even-if-the-internet-disconnects-cab5a3b0abae
- url
- https://medium.com/@codefarm0/system-design-interview-how-would-you-let-users-upload-huge-files-even-if-the-internet-disconnects-cab5a3b0abae
- canonical_url
- https://medium.com/@codefarm0/system-design-interview-how-would-you-let-users-upload-huge-files-even-if-the-internet-disconnects-cab5a3b0abae
- author_url
- https://medium.com/@codefarm0
- status
- ok
- fetched_at
- 2026-07-21 04:28:33