← Back to Blog
Security11 min read

File Upload Vulnerabilities: The Complete Pentester Guide (2026)

October 2, 2026

File upload functionality sits at the intersection of three difficult engineering problems — content type validation, filesystem storage, and content rendering. Each is independently tricky; combined they produce one of the most consistently exploitable attack surfaces in modern web applications. OWASP has listed unrestricted file upload as a severe vulnerability class since 2004 and the finding still appears routinely in pentest reports two decades later. This document covers the modern exploitation workflow, the specific bypass techniques that still work against real applications, and the defensive architecture that reliably blocks them.

Why File Upload Keeps Breaking

The underlying difficulty is that validating "this file is safe" requires knowing what the file will do when it reaches each system that handles it — the web server, the application framework, the storage backend, the image processor, the virus scanner, the download endpoint, and ultimately the client that renders it. Each of these systems parses the file differently, and attackers exploit the differences. A file that validates as a benign PNG to one system can execute as PHP to another.

Validation strategies that focus on a single property — extension, content-type header, magic bytes — all fail against adversarial inputs because the attacker controls the input. The validation needs to be defence in depth: multiple independent checks that each eliminate a class of attack, plus storage and rendering architecture that reduces the impact of validation failures.

The Impact Spectrum

File upload vulnerabilities range in impact from information disclosure to remote code execution depending on where the uploaded file lands and how it is processed. The highest-impact scenario is file upload to a web-accessible directory where the server executes the file as code — classical PHP/ASP/JSP shell upload. This produces immediate remote code execution with the privileges of the web server process.

Below RCE, upload-to-storage with server-side processing by vulnerable libraries produces impact via the library (ImageMagick RCE via crafted PNG, PDF renderer RCE via crafted PDF, archive extraction RCE via zip slip). Upload-to-rendered-content produces XSS when the uploaded file is served with content types the browser executes. Upload-to-shared-storage produces information disclosure when path traversal in filenames escapes the intended directory. Each tier has its own exploitation workflow and its own defensive requirements.

Extension-Based Bypass Techniques

The weakest form of file upload validation is extension whitelisting or blacklisting based on the filename only. Bypass techniques against this layer are old but still work against surprisingly many applications.

Double extensions: file.php.jpg or file.jpg.php depending on the server behaviour. Apache with mod_php often executes .php.jpg as PHP because the AddHandler directive matches on any extension in the filename. The fix is to use SetHandler and match specifically on the terminal extension.

Case variation: file.PHP, file.PhP, file.pHp. Blacklists that check against lowercase .php miss these on case-insensitive filesystems. Validation must either lowercase before checking or match case-insensitively.

Alternative extensions that resolve to the same handler: .php3, .php4, .php5, .php7, .phtml, .phar for PHP; .asp, .aspx, .ashx, .asmx, .cshtml for ASP.NET; .jsp, .jspx, .jhtml, .jspf for Java. Each handler has multiple extensions configured by default on many servers. Blacklists that cover only the primary extension miss the alternatives.

Null byte injection: file.php%00.jpg. On older PHP versions and some C-based image processors, the null byte terminates the string at the filesystem layer while the extension check sees the trailing .jpg. Modern PHP has fixed this but legacy code paths and non-PHP applications still occasionally exhibit the behaviour.

Trailing characters: file.php. (trailing dot), file.php%20 (trailing space), file.php:: (NTFS alternate data stream). Windows filesystems in particular strip trailing whitespace and dots at write time, so file.php. becomes file.php on disk while the pre-write extension check saw a different string.

Content-Type and MIME Bypass

Content-Type validation based on the client-supplied header is trivially bypassed — the attacker controls the request and can set Content-Type: image/jpeg on any payload. This is the Burp-intercepted-request bypass that still appears in CTF scenarios because it still works against real applications.

Server-side MIME detection via libmagic or the file command adds a stronger check but is still bypassable. The common technique is content-type confusion: craft a file that begins with valid magic bytes for an allowed type but continues with malicious content. A PNG with PHP code appended after the IEND chunk passes libmagic's PNG detection but executes as PHP when the server processes it. GIF89a; prepended to a PHP shell produces the same result. Image parsers typically do not reject trailing garbage, and the magic-bytes check does not inspect past the signature.

Polyglot files — valid as multiple formats simultaneously — are the sophisticated version of this attack. A GIF/PHP polyglot renders as a GIF when the browser requests it and executes as PHP when the server interprets it. GIFAR polyglots (GIF + JAR) were historically used to execute Java applets served from an image URL. The modern equivalents include JPG/PHP, PNG/HTML, and PDF/HTML polyglots that each bypass single-format validation while producing attacker-controlled behaviour in the eventual handler.

Server Configuration Bypass

Even when the application validates uploads strictly, server configuration bypasses can produce the same outcome. The .htaccess upload bypass works when the application allows .htaccess files to be uploaded to the upload directory — the uploaded .htaccess adds new handlers that reinterpret innocent-looking files as PHP. Similar bypasses exist for web.config in IIS environments, where uploaded web.config overrides server behaviour at the directory level.

IIS short filename (8.3) attacks exploit the fact that NTFS maintains legacy 8.3 filenames alongside long filenames. An uploaded shell.aspx file is accessible as SHELL~1.ASP via the short name, which may bypass extension-filter rules that only check the long name.

Server-side request routing based on URL path extension can also be exploited. A file uploaded as image.jpg but accessed via image.jpg/x.php may route through PHP interpretation on servers with path-info enabled. The exact behaviour depends on the server configuration but the pattern has produced many historical RCEs.

Path Traversal in Filenames

Filename handling independently produces path traversal vulnerabilities. The uploaded filename, if used in filesystem operations without sanitisation, can contain ../ sequences that escape the intended upload directory. The attacker uploads a file named ../../../var/www/html/shell.php and achieves file placement outside the sandbox.

Zip slip is the archive-extraction variant. The uploaded archive contains entries with ../ in their internal paths; naive extraction code writes those entries to the computed destination, which lands outside the extraction directory. Zip slip was identified as a widespread class in 2018 and still appears in applications that use default archive extraction libraries without path validation.

Image Processing Exploits

Applications that process uploaded images server-side introduce an entirely separate attack surface. ImageMagick vulnerabilities have been a consistent source of RCE for a decade — ImageTragick (CVE-2016-3714) was the most famous, but subsequent CVEs appear every few months. The exploitation pattern is: upload a crafted image file, trigger server-side processing by requesting a thumbnail or resize, achieve RCE through the processing library.

Defensive measures include disabling the vulnerable coders in ImageMagick's policy.xml (specifically EPHEMERAL, URL, MVG, MSL, TEXT, SHOW, WIN, PLT), using Pillow or other alternative libraries where the attack surface is smaller, or sandboxing the image processing in a separate process with no network access and minimal filesystem access.

SVG upload is a specific case worth separate treatment. SVG files are XML and can contain JavaScript (via onload handlers), XXE payloads, and references to external resources. SVGs that reach the client browser execute their embedded JavaScript as part of the hosting page's origin — producing stored XSS on any application that renders uploaded SVGs inline. The defence is to serve SVGs with Content-Type: image/svg+xml and Content-Security-Policy that blocks script execution, or to rasterise SVGs to PNG server-side before serving.

The CSV Injection Variant

CSV upload with downstream Excel opening produces CSV injection — formula execution when a cell begins with =, +, -, or @. The injected formula can execute via DDE, exfiltrate data via HYPERLINK, or trigger external process launch via WEBSERVICE. The vulnerability is in the Excel-side rendering rather than the server, but the server is the attack delivery vector.

Defensive mitigation requires prefixing cells that begin with the trigger characters with a single quote during export, or validating uploaded CSVs for the pattern and rejecting them. This is particularly important for applications that re-export user-uploaded data to other users, since the attacker's CSV ends up in victim users' Excel.

Detection Workflow for Pentesters

A systematic file upload testing workflow has four phases. First, map the upload surface: identify every endpoint that accepts files, including profile pictures, document uploads, support ticket attachments, and import features. Each endpoint has its own validation logic and should be tested independently.

Second, baseline the validation. Upload a known-good file and observe the response, the storage location (if observable), and any processing that occurs. Note the Content-Type the server records, the filename it assigns, and any URL where the file becomes accessible.

Third, probe the validation layers systematically. Vary the extension (double, case, alternative, null byte, trailing character). Vary the Content-Type header independently of the content. Craft polyglot files that are valid as the allowed type and malicious as a different handler. Attempt path traversal in the filename.

Fourth, probe the processing layer. If the server processes uploaded images, test the processor's vulnerable coders. If the server extracts uploaded archives, test zip slip. If the server reads uploaded CSVs, test CSV injection. Each processing step is a separate attack surface regardless of the upload validation.

Defensive Architecture

The reliable defensive architecture has six components, each eliminating a class of attack. Extension validation uses a strict whitelist of allowed extensions matched against the final extension of the filename, case-insensitive, with no exceptions for alternative handlers. Content-Type validation uses server-side detection (libmagic or equivalent) rather than client-supplied headers.

Storage architecture places uploaded files in a separate, non-executable directory served by a dedicated subdomain or path. The web server is explicitly configured to not execute any content from the upload directory regardless of extension. For high-security applications, uploaded files are stored in object storage (S3, Azure Blob) rather than on the application server filesystem, which eliminates the local-execution attack entirely.

Filename handling generates server-side filenames (typically UUID-based) rather than preserving the user-supplied filename. The original filename is stored separately in the database for display purposes but never participates in filesystem operations. This eliminates path traversal in filenames completely.

Content processing runs in isolated, resource-limited environments. Image processing happens in a sandboxed service with no network access, limited filesystem access, and strict memory limits. Archive extraction uses libraries with path-traversal protection enabled. Document rendering uses sandboxed renderers (headless browsers in containers, dedicated PDF renderers with no shell access).

Serving architecture uses Content-Type headers derived from server-side detection rather than client-supplied headers, X-Content-Type-Options: nosniff to prevent browser MIME sniffing, and Content-Disposition: attachment where feasible to prevent inline rendering. Content-Security-Policy on the hosting application blocks script execution from the upload origin.

Finally, antivirus scanning of uploaded files catches known-malicious payloads before storage. AV catches a different class of attack than the structural defences above — it blocks commodity malware that users attempt to share, which complements the exploitation-focused controls. Running both ClamAV and a cloud-based scanner (via API) produces better detection than either alone.

What Pentest Reports Should Flag

File upload findings should distinguish between the specific bypass that succeeded and the broader architectural weakness. A finding that reports "extension validation bypass via double extension" is less actionable than one that reports "upload validation relies solely on filename extension, bypassable through multiple techniques, and the storage architecture executes uploaded content as code — recommended fix is defence in depth across extension validation, Content-Type validation, storage isolation, and server configuration."

The remediation guidance should also address the specific technology stack. For PHP applications, the fix includes SetHandler configuration, extension whitelist, and non-executable upload directory. For Node.js, the fix often focuses on filename handling and sandboxed processing because the execution-as-code vector is weaker. For Java/.NET, the fix emphasises deserialisation safety in file processing and strict Content-Type configuration. Generic file-upload remediation advice without stack-specific detail typically does not get implemented.

Stop finding vulnerabilities manually

TigerStrike uses AI agents to continuously discover, validate, and exploit vulnerabilities across your applications — so your team can focus on fixing what matters.