how to validate file uploads securely
A step-by-step skill for locking down file uploads: type and size checks, safe filenames, storage outside the web root, and serving files without executing them. Use when an agent reviews an upload feature, builds avatar or document upload, or audits where uploaded files land. Triggers: 'secure file upload', 'validate file uploads', 'upload validation checklist', 'is this upload handler safe'. Not for: image-only CDN pipelines, virus scanning products, or general form validation.
how to validate file uploads securely
TL;DR
Treat every uploaded file as hostile: check the actual content type, not the extension; cap the size; rename it to something random; store it outside the web root; and serve it back so it can never execute. Most upload bugs are one missing check in that chain.
how to validate file uploads securelyUse this when
- You are building or reviewing avatar, document, or media upload
- A PR adds multipart handling or a new accepted file type
- Uploaded files are served back to other users
- An audit asks where user-supplied bytes end up on disk
- An agent is checking an upload endpoint for traversal or execution risk
Not for this skill when
- Files come from a trusted internal pipeline, not users
- You need enterprise malware sandboxing (product territory)
- The question is about resumable-upload UX or progress bars
- You are validating text form fields, not files
Steps
1. Enforce an allowlist of types, then verify content
Decide which types you accept (images, PDFs, whatever the feature needs) and check the file's magic bytes server-side. Extensions and the client-sent content type are just hints.
file --mime-type [HOME]/...Expected: the reported MIME type is on your allowlist. A file named photo.png that reports as a script or archive gets rejected, and the rejection is logged.
2. Cap the size before you read the file
Set a max size in the web server or framework config so a huge upload is refused during transfer, not after you have buffered it into memory.
grep -rn "client_max_body_size\|MAX_CONTENT_LENGTH\|maxFileSize" [HOME]/...Expected: a limit exists and is small enough for the feature (a few MB for avatars, not hundreds). Oversized uploads get a clean 413, not a crashed worker.
3. Rename with a random name and drop the original path
Generate the stored filename yourself with a UUID and keep the original name only as display metadata. Never let the upload path influence where the file lands.
python3 -c "import uuid; print(uuid.uuid4().hex + '.png')"Expected: stored names are random and extension comes from your verified type, not the upload. Path traversal sequences in the original filename go nowhere.
4. Store outside the web root and serve deliberately
Put uploads in a directory the web server never executes, and serve them through an endpoint that sets a safe content type and a download disposition for anything active.
Expected: requesting the stored file never runs it as a page or script. Uploading an HTML file and opening its URL shows inert content or triggers a download, never a rendered page in your origin.
5. Scan and quarantine when the stakes are high
If files are shared between users, run an antivirus scan on upload and hold the file until it passes. This is defense in depth, not a substitute for steps 1 to 4.
Expected: the pipeline has a scan stage with a pass or quarantine outcome, and quarantined files are never served.
6. Test the nasty cases
Upload a file with a double extension, a null byte in the name, an oversized body, a mismatched magic number, and an SVG containing script. Confirm each is rejected or neutralized.
Expected: all five probes are handled safely, and the outcomes are written down as regression tests.
Variant: avatar and image uploads
Re-encode images through a processing library on upload. Re-encoding strips embedded scripts, normalizes the format, and gives you thumbnails for free. Reject anything the decoder chokes on.
Variant: document uploads shared between users
Serve with content-disposition attachment so browsers download instead of rendering, and consider a viewer that converts to a safe preview rather than serving the original bytes.
Variant: direct-to-storage uploads
When the browser uploads straight to object storage, sign the policy server-side with tight constraints on size, type, and key prefix, and validate again when your backend first touches the file.
Variant: CSV and spreadsheet imports
CSVs get formula injection: cells starting with =, +, -, or @ can execute in spreadsheet apps. Prefix or strip those characters on import when the file will be opened in Excel or Sheets.
Why this happens
Upload handlers sit at the exact boundary where outside bytes become inside files, and every shortcut (trusting the extension, keeping the original filename, serving from a web-accessible path) hands the attacker control over one more link. The chain works because each step removes one attacker-controlled variable: type, size, name, location, and execution context.
Edge cases and pitfalls
- Polyglot files are valid as two formats at once; re-encoding (images) or conversion (documents) beats detection.
- ZIP bombs and decompression bombs defeat size caps on the compressed file; cap the decompressed size too.
- Race conditions between the type check and the move can be exploited; check and store atomically where you can.
- Old files uploaded before the fix stay dangerous; backfill or quarantine the existing bucket.
- Logging original filenames is fine, but render them as text in any admin UI, never as HTML.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_0ilaJiHnrlznq6yT9NIIQQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.