commit c844e4ea5fb22ffd8f4f7568786d2273eb8f69ed Author: utkusen Date: Mon Mar 30 15:22:07 2026 +0100 first commit diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..e43b0f9 --- /dev/null +++ b/.gitignore @@ -0,0 +1 @@ +.DS_Store diff --git a/README.md b/README.md new file mode 100644 index 0000000..184e840 --- /dev/null +++ b/README.md @@ -0,0 +1,71 @@ +# LLM SAST Skills + +A collection of agent skills that turn your LLM coding assistant into a fully functional SAST scanner to find vulnerabilities in your codebase. Works natively with Claude Code, Codex, Opencode, Cursor and any other assistant that supports agent skills. No third-party tools required. + +Claude Code with Opus model is recommended. But if the cost is a concern, use any IDE and model you trust. + +![Screenshot from Opencode](opencode-screenshot.png) + +## How It Works + +`CLAUDE.md` (for Claude Code) or `AGENTS.md` (for Opencode and other IDEs) orchestrates the entire assessment workflow automatically. The assessment runs in three steps: + +1. **Codebase Analysis** -- The `sast-analysis` skill maps the technology stack, architecture, entry points, data flows, and trust boundaries. It writes its findings to `sast/architecture.md`. + +2. **Vulnerability Detection (parallel)** -- All 13 vulnerability detection skills run in parallel as subagents. Each skill follows a two-phase approach: first a recon/discovery phase to find candidate sections, then a verification phase to confirm exploitability. Results are written to `sast/*-results.md`. + +3. **Report Generation** -- The `sast-report` skill consolidates all findings into a single `sast/final-report.md`, ranked by severity with full remediation guidance and dynamic test instructions. + +## What It Detects + +| Skill | Vulnerability Class | +|---|---| +| sast-analysis | Codebase reconnaissance, architecture mapping, threat modeling | +| sast-sqli | SQL Injection | +| sast-graphql | GraphQL injection | +| sast-xss | Cross-Site Scripting (XSS) | +| sast-rce | Remote Code Execution (command injection, eval, unsafe deserialization) | +| sast-ssrf | Server-Side Request Forgery | +| sast-idor | Insecure Direct Object Reference | +| sast-xxe | XML External Entity | +| sast-ssti | Server-Side Template Injection | +| sast-jwt | Insecure JWT implementations | +| sast-missingauth | Missing authentication and broken function-level authorization | +| sast-pathtraversal | Path / directory traversal | +| sast-fileupload | Insecure file upload | +| sast-businesslogic | Business logic flaws (price manipulation, workflow bypass, race conditions, etc.) | +| sast-report | Consolidated final report ranked by severity | + + +## Installation + +Copy your project into the `sast-files` folder, then open `sast-files` as your workspace in your AI coding assistant. + +```bash +cp -r /path/to/your/project sast-files/ +``` + +> **Note:** If your project already contains a `CLAUDE.md` or `AGENTS.md` file, remove it before running the assessment — otherwise it will conflict with the orchestration file provided by this toolkit. + + +## Usage + +After copying the files, open your project in your AI coding assistant and ask: + +> Run vulnerability scan + +or + +> Find vulnerabilities in this codebase + +The entry point file (`CLAUDE.md` or `AGENTS.md`) orchestrates the full workflow automatically. It will skip any steps whose output files already exist, so you can safely re-run it after fixing issues. + +## Output + +All output is written to a `sast/` folder in your project root: + +| File | Description | +|---|---| +| `sast/architecture.md` | Technology stack, architecture, entry points, data flows | +| `sast/*-results.md` | Per-vulnerability-class findings with proof and remediation | +| `sast/final-report.md` | Consolidated report ranked by severity | diff --git a/opencode-screenshot.png b/opencode-screenshot.png new file mode 100644 index 0000000..997a996 Binary files /dev/null and b/opencode-screenshot.png differ diff --git a/sast-files/.agents/skills/sast-analysis/SKILL.md b/sast-files/.agents/skills/sast-analysis/SKILL.md new file mode 100644 index 0000000..6b6f1ab --- /dev/null +++ b/sast-files/.agents/skills/sast-analysis/SKILL.md @@ -0,0 +1,91 @@ +--- +name: sast-analysis +description: >- + Perform codebase analysis and architecture mapping as the first phase of a + security assessment. Explores the tech stack, frameworks, entry points, data + flows, and trust boundaries. Outputs sast/architecture.md. Run this before any + vulnerability detection skill. Use when asked to analyze a codebase for + security or when sast/architecture.md does not yet exist. +--- + +# Codebase Analysis + +You are performing the first phase of a security assessment. Your goal is to deeply understand the codebase. You are NOT looking for specific vulnerabilities yet. This is pure reconnaissance. + +Create a `sast/` folder in the project root (if it doesn't already exist). This phase produces one output file inside it: + +`sast/architecture.md` — technology stack, architecture, entry points, data flows + +## Phase 1: Technology Reconnaissance + +Explore the codebase and identify: + +- **Languages**: All programming languages used and their versions if specified +- **Frameworks**: Web frameworks, ORM layers, template engines, task queues +- **Package managers & dependencies**: Lock files, dependency manifests (package.json, requirements.txt, go.mod, Gemfile, pom.xml, etc.) +- **Infrastructure hints**: Dockerfiles, docker-compose, Kubernetes manifests, Terraform, CI/CD configs +- **Databases**: SQL, NoSQL, cache layers, message brokers — look at connection strings, ORM models, migration files +- **Authentication & authorization**: Auth libraries, middleware, session configs, OAuth/OIDC providers, JWT usage, API key patterns +- **External integrations**: Third-party APIs, payment processors, email services, cloud SDKs, webhook handlers +- **Entry points**: HTTP routes, GraphQL schemas, gRPC service definitions, CLI commands, WebSocket handlers, scheduled jobs, message consumers + +Start by reading dependency manifests, project configs, and directory structure. Then drill into source code to confirm findings. + +## Phase 2: Architecture Mapping + +Based on Phase 1, build a mental model of: + +1. **Service boundaries**: Is this a monolith or microservices? What talks to what? +2. **Data flow**: How does user input enter the system, get processed, get stored, and get returned? +3. **Trust boundaries**: Where does the system transition between trusted and untrusted contexts? (e.g., user input -> backend, backend -> database, service -> service, server -> client) +4. **Privilege levels**: What roles/permissions exist? How are they enforced? Is there an admin panel? +5. **Sensitive data inventory**: PII, credentials, tokens, financial data, health records — where is each stored and how does it move? + +**Write the results of Phase 1 and Phase 2 to `sast/architecture.md`.** Use this format: + +```markdown +# Architecture: [Project Name] + +## Technology Stack + +| Category | Details | +|---|---| +| Languages | ... | +| Frameworks | ... | +| Databases | ... | +| Auth mechanism | ... | +| Infrastructure | ... | +| External services | ... | + +## Architecture Overview + +[Describe the architecture: monolith vs microservices, how components interact, +main modules and their responsibilities] + +## Data Flow + +[Trace how user input enters the system, gets processed, stored, and returned. +Cover the primary flows (e.g., registration, login, core business actions).] + +## Entry Points + +| Entry Point | Type | Auth Required | Description | +|---|---|---|---| +| ... | HTTP/GraphQL/WS/etc. | Yes/No | ... | + +## Trust Boundaries + +[List each trust boundary and what crosses it] + +## Sensitive Data Inventory + +| Data Type | Where Stored | How Accessed | Protection | +|---|---|---|---| +| ... | ... | ... | ... | +``` + +## Important Reminders + +- Do NOT report specific vulnerabilities (like "line 42 has SQL injection"). That comes in later phases. +- Be thorough in exploration. Read actual source code, not just config files. Look at how auth middleware is applied, how queries are built, how file uploads are handled. +- If the codebase is large, prioritize security-sensitive areas: auth, payment, data access, file handling, admin functionality. diff --git a/sast-files/.agents/skills/sast-businesslogic/SKILL.md b/sast-files/.agents/skills/sast-businesslogic/SKILL.md new file mode 100644 index 0000000..d7f331f --- /dev/null +++ b/sast-files/.agents/skills/sast-businesslogic/SKILL.md @@ -0,0 +1,326 @@ +--- +name: sast-businesslogic +description: >- + Detect business logic vulnerabilities in a codebase using a two-phase + approach: first perform threat modeling by analyzing the application's + domain and generating specific attack scenarios (price manipulation, + workflow bypass, limit violations, race conditions, reward abuse, etc.), + then verify whether those threats are exploitable by checking for missing + validations and enforcement. Requires sast/architecture.md (run + sast-analysis first). Outputs findings to sast/businesslogic-results.md. + Use when asked to find business logic, logic flaws, or abuse-of-function bugs. +--- + +# Business Logic Vulnerability Detection + +You are performing a focused security assessment to find business logic vulnerabilities in a codebase. This skill uses a two-phase approach with subagents: **threat modeling** (understand the domain and generate attack scenarios) then **verify** (check whether those attack scenarios are exploitable). + +**Prerequisites**: `sast/architecture.md` must exist. Run the analysis skill first if it doesn't. + +--- + +## What are Business Logic Vulnerabilities + +Business logic vulnerabilities arise when an application's intended workflow, rules, or constraints can be manipulated to produce unintended outcomes — without exploiting technical flaws like injection or memory corruption. The attacker operates within the application's own features but uses them in ways the developers did not anticipate. + +The core pattern: *the application accepts input that is syntactically valid and passes authentication/authorization, but violates a business rule that was never enforced in code.* + +### What Business Logic Vulnerabilities ARE + +- Submitting a negative quantity to a purchase endpoint, receiving a credit instead of a charge +- Applying the same one-time discount coupon multiple times in parallel requests +- Skipping the payment step in a multi-step checkout by replaying a later step's request +- Posting a rating of 9999 to a movie rating endpoint that should cap ratings at 5 +- Transferring a negative amount to move money from the recipient to the sender +- Redeeming a referral bonus by referring yourself with a second account +- Re-using a single-use reset token or voucher that was never invalidated +- Purchasing an item that is out of stock due to a race condition between inventory check and reservation +- Accessing a premium subscription feature after downgrading to a free plan +- Winning an auction by retracting a high bid after others have been eliminated + +### What Business Logic Vulnerabilities are NOT + +Do not flag these as business logic issues: + +- **SQL injection, XSS, RCE, XXE, SSRF, SSTI**: These are injection/technical flaws — separate skills cover them +- **Missing authentication**: Endpoint requires no login at all → that's "Unauthenticated Access" +- **IDOR**: Accessing another user's resource by changing an ID → that's a separate access-control class +- **Brute-force / rate limiting**: Generic rate-limit bypass on login → that's not a business logic flaw unless it enables specific business rule circumvention + +--- + +## Business Logic Attack Categories + +Use these categories to guide threat modeling. Not all categories apply to every application — identify which ones are relevant based on the architecture summary. + +### 1. Price & Payment Manipulation +- Negative prices or zero prices on purchase endpoints +- Arbitrary price override in request body (mass assignment of price field) +- Currency or unit confusion (e.g., cents vs. dollars) +- Floating-point precision abuse in monetary arithmetic +- Applying discounts that reduce total below zero + +### 2. Quantity & Numeric Limit Violations +- Negative quantities (ordering −5 items to receive a credit) +- Quantities exceeding per-user or per-order limits +- Integer overflow/underflow in quantity or balance calculations +- Out-of-range values for bounded fields (ratings, scores, percentages) + +### 3. Workflow & Multi-Step Process Bypass +- Skipping mandatory steps in a sequential process (payment, email verification, ID check) +- Replaying a completion token from a previous successful flow to bypass steps +- Direct-access to a later-stage endpoint without completing earlier stages +- Submitting a terminal state transition without going through intermediate states (state machine violations) + +### 4. Coupon, Discount & Voucher Abuse +- Applying the same coupon multiple times (single-use not enforced) +- Stacking discounts that were not intended to be combined +- Using an expired coupon or voucher +- Generating or guessing valid coupon codes + +### 5. Race Conditions & Concurrency Abuse +- Double-spending: sending two concurrent purchase requests to consume a balance once +- Concurrent coupon redemption draining credit beyond allowed amount +- TOCTOU (time-of-check / time-of-use) on inventory: check passes for both requests, both reservations succeed +- Parallel withdrawal/transfer requests exceeding account balance + +### 6. Refund & Chargeback Abuse +- Requesting a refund after the digital good has been consumed or downloaded +- Partial refund on an already-partially-refunded order +- Refund without returning physical item (if logic is not enforced server-side) + +### 7. Reward, Referral & Loyalty Abuse +- Self-referral using a second account to earn a referral bonus +- Earning signup bonuses multiple times across multiple accounts +- Loyalty point farming through artificial activity +- Sharing or transferring non-transferable rewards + +### 8. Subscription & Entitlement Bypass +- Accessing paid/premium features after downgrading or cancelling +- Trial period abuse (repeatedly creating new accounts for trial access) +- Feature flag or plan check performed only at subscription creation, not at feature access time +- Entitlement cached at session start and not re-evaluated after plan change + +### 9. Auction & Bidding Logic +- Retracting a winning bid after competing bids have been rejected +- Shill bidding: artificially inflating price with controlled accounts +- Bypass of reserve price enforcement +- Bid manipulation via concurrent requests + +### 10. Inventory & Stock Logic +- Purchasing out-of-stock items due to missing stock validation +- Reserving more stock than available via concurrent requests +- Negative inventory resulting from refund-without-restock logic +- Phantom inventory: item appears available but cannot be fulfilled + +### 11. Time & Date Logic +- Using time-limited offers after expiration (expiry checked client-side or weakly server-side) +- Backdating transactions or bookings +- Exploiting "grace period" logic to extend benefits indefinitely +- System clock manipulation if server trusts client-supplied timestamps + +### 12. Transfer & Balance Logic +- Transferring a negative amount (sender receives money from recipient) +- Self-transfer to exploit bonus or fee logic +- Transferring more than the available balance due to missing server-side check +- Rounding errors exploited across many micro-transactions + +--- + +## Execution + +This skill runs in two phases using subagents. Pass the contents of `sast/architecture.md` to both subagents as context. + +### Phase 1: Threat Modeling — Domain Analysis & Attack Scenario Generation + +Launch a subagent with the following instructions: + +> **Goal**: Analyze the codebase to understand its business domain and generate a concrete, prioritized list of business logic attack scenarios specific to this application. Write results to `sast/businesslogic-threats.md`. +> +> **Context**: You will be given the project's architecture summary. Use it to understand what the application does, what features it has, and what business rules it is supposed to enforce. Focus entirely on understanding the domain — do not verify vulnerabilities yet. +> +> **Step 1 — Identify the business domain and features**: +> +> Read `sast/architecture.md` and then explore the codebase to answer: +> - What does this application do? (e-commerce, marketplace, SaaS, social platform, fintech, gaming, booking, etc.) +> - What financial or transactional features exist? (payments, subscriptions, credits, tokens, wallets, invoices, refunds) +> - What quantitative limits or rules exist? (ratings, scores, quantities, usage limits, quotas) +> - What multi-step workflows exist? (checkout, onboarding, KYC, booking, auctions) +> - What promotional or reward features exist? (coupons, referrals, loyalty points, bonuses, vouchers) +> - What role or tier distinctions exist? (free vs. paid, user vs. premium, trial vs. full) +> - What inventory or capacity constraints exist? (stock, seats, slots, bandwidth) +> +> To discover features, search for: +> - Route/endpoint definitions and their names +> - Model/entity names (Order, Payment, Subscription, Coupon, Wallet, Bid, etc.) +> - Business-rule-related field names (price, quantity, balance, rating, score, limit, quota, expiry, status) +> - Validation logic or constraint-related code +> +> **Step 2 — Generate attack scenarios**: +> +> For each relevant business domain area found, generate specific attack scenarios. Each scenario must be: +> - **Specific to this codebase** — name the actual endpoint, model, or feature involved +> - **Actionable** — describe exactly what an attacker would send/do +> - **Grounded** — reference the code or data model that makes this scenario plausible +> +> Use the attack categories below as a checklist. Only include categories that are relevant to this application: +> +> - **Price/payment manipulation**: Can a user send an arbitrary price in the request? Is price trusted from client? +> - **Quantity/value out of range**: Can a user send negative quantities, zero, or values exceeding defined limits? +> - **Workflow bypass**: Can a user skip a mandatory step in a multi-step process? +> - **Coupon/discount abuse**: Can a coupon be used multiple times or after expiration? +> - **Race conditions**: Are there check-then-act patterns on shared resources (inventory, balance, coupon usage)? +> - **Refund abuse**: Can a refund be requested after the product is consumed? +> - **Reward/referral abuse**: Can referral or signup bonuses be farmed? +> - **Entitlement bypass**: Are premium features checked at access time or only at subscription time? +> - **Transfer/balance logic**: Can negative transfers or self-transfers be made? +> - **Time/date logic**: Are time-limited offers enforced server-side? +> - **Inventory logic**: Is stock validated atomically before reservation? +> +> **Output format** — write to `sast/businesslogic-threats.md`: +> +> ```markdown +> # Business Logic Threat Model: [Project Name] +> +> ## Application Domain +> [2–3 sentence summary of what the application does and its key business features] +> +> ## Business Features Identified +> - [Feature 1]: [brief description, relevant models/endpoints] +> - [Feature 2]: ... +> +> ## Attack Scenarios +> +> ### Scenario 1: [Short title, e.g. "Negative quantity purchase for credit"] +> - **Category**: [e.g. Quantity & Numeric Limit Violations] +> - **Target**: [Endpoint or feature, e.g. `POST /api/orders`] +> - **Description**: [What an attacker would do and what outcome they expect] +> - **Relevant code**: [File and line range where the relevant logic lives] +> - **Business rule that should be enforced**: [What the application is supposed to do] +> - **Risk level**: [High / Medium / Low] +> +> ### Scenario 2: ... +> +> ## Categories Not Applicable +> [List any categories from the checklist that are not relevant to this application and why] +> ``` + +### Phase 2: Verify — Check Whether Scenarios Are Exploitable + +Launch a second subagent **after Phase 1 completes** with the following instructions: + +> **Goal**: For each attack scenario in `sast/businesslogic-threats.md`, determine whether the business rule is properly enforced in code or whether the attack is exploitable. Write final results to `sast/businesslogic-results.md`. +> +> **Context**: You will be given the project's architecture summary and the threat model. Use the architecture summary to understand validation patterns, ORM usage, and where business rules are typically enforced. +> +> **For each scenario, perform the following checks**: +> +> **1. Is the business rule enforced server-side?** +> - Is the constraint validated in the backend handler, service layer, or ORM/database? +> - Or is it only validated client-side (frontend form validation, JavaScript min/max attributes)? +> - Client-side-only validation = exploitable. +> +> **2. Is the validation complete and covers all edge cases?** +> - Does it check for negative values where applicable? +> - Does it check upper bounds, not just lower bounds? +> - Does it handle concurrent requests (is the check atomic, or is there a TOCTOU window)? +> - Does it re-validate at the point of use, not just at an earlier step? +> +> **3. For workflow bypass scenarios**: +> - Does each step verify that previous required steps were completed? +> - Are step completion flags stored server-side (not just in a cookie or session that can be replayed)? +> - Can a terminal endpoint be called directly without going through earlier steps? +> +> **4. For coupon/voucher scenarios**: +> - Is the coupon marked as used atomically with the transaction (in the same DB transaction)? +> - Is concurrent redemption protected (SELECT FOR UPDATE, optimistic locking, atomic compare-and-swap)? +> - Is the expiry date checked server-side at redemption time? +> +> **5. For race condition scenarios**: +> - Is stock/balance check and decrement done atomically (in a single DB transaction or with row-level locking)? +> - Is there any idempotency key or deduplication logic to prevent duplicate concurrent requests? +> +> **6. For entitlement/subscription scenarios**: +> - Is the user's current plan/tier checked at the point of feature access? +> - Or is it cached at login/session start and never re-evaluated? +> +> **7. For transfer/balance scenarios**: +> - Is there a server-side check that the transfer amount is positive? +> - Is there a server-side check that the sender has sufficient balance? +> - Are these checks done within a database transaction to prevent race conditions? +> +> **Classification**: +> - **Exploitable**: The business rule is absent, bypassable, or only enforced client-side. +> - **Likely Exploitable**: The rule exists but has gaps (race condition window, missing edge case, bypassable condition). +> - **Not Exploitable**: Proper server-side enforcement exists and covers edge cases. +> - **Needs Manual Review**: Cannot determine with confidence (complex logic, external service dependency, etc.). +> +> **Output format** — write to `sast/businesslogic-results.md`: +> +> ```markdown +> # Business Logic Analysis Results: [Project Name] +> +> ## Executive Summary +> - Scenarios analyzed: [N] +> - Exploitable: [N] +> - Likely Exploitable: [N] +> - Not Exploitable: [N] +> - Needs Manual Review: [N] +> +> ## Findings +> +> ### [EXPLOITABLE] Scenario title +> - **Category**: [Attack category] +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint**: `METHOD /path` +> - **Business Rule Violated**: [What rule the application should enforce] +> - **Issue**: [Clear description of what validation is missing or broken] +> - **Impact**: [What an attacker can achieve — free goods, financial loss, unfair advantage, etc.] +> - **Proof**: [Show the code path demonstrating the missing enforcement] +> - **Remediation**: [Specific fix for this scenario] +> - **Dynamic Test**: +> ``` +> [Step-by-step instructions or curl commands to confirm the finding on the live app. +> Include exact HTTP method, endpoint, headers, and request body. +> Describe what response or side effect confirms the vulnerability.] +> ``` +> +> ### [LIKELY EXPLOITABLE] Scenario title +> - **Category**: [Attack category] +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint**: `METHOD /path` +> - **Business Rule Violated**: [What rule should be enforced] +> - **Issue**: [What enforcement gap or race condition exists] +> - **Concern**: [Why this is likely exploitable despite partial enforcement] +> - **Proof**: [Show the code path with the weak/partial check] +> - **Remediation**: [Specific fix] +> - **Dynamic Test**: +> ``` +> [Step-by-step instructions or curl commands, e.g. two concurrent requests, to confirm.] +> ``` +> +> ### [NOT EXPLOITABLE] Scenario title +> - **Category**: [Attack category] +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Business Rule**: [What the application is supposed to enforce] +> - **Protection**: [How it is enforced — server-side validation, DB constraint, atomic transaction, etc.] +> +> ### [NEEDS MANUAL REVIEW] Scenario title +> - **Category**: [Attack category] +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Uncertainty**: [Why automated analysis couldn't determine the status] +> - **Suggestion**: [What to examine manually or test dynamically] +> ``` + +--- + +## Important Reminders + +- Read `sast/architecture.md` and pass its content to both subagents as context. +- Phase 2 must run **after** Phase 1 completes — it depends on the threat model output. +- Focus strictly on **business logic flaws** — do not flag injection bugs, auth bypass, or IDOR issues here. +- Threat modeling in Phase 1 should be **application-specific**: generic scenarios not grounded in the actual codebase are not useful. +- Server-side validation is the only valid protection. Client-side validation, frontend form constraints, and API documentation that says "must be positive" are not security controls. +- Race conditions on financial operations are high-severity even if they appear to require exact timing — automated tools (Turbo Intruder, concurrent curl) make them trivial to exploit. +- When in doubt, classify as "Needs Manual Review" rather than "Not Exploitable". False negatives in a security assessment are worse than false positives. +- Pay attention to ORM and database-level constraints (CHECK constraints, unique indexes, transactions with locking) — these can provide enforcement that is not visible in application code alone. diff --git a/sast-files/.agents/skills/sast-fileupload/SKILL.md b/sast-files/.agents/skills/sast-fileupload/SKILL.md new file mode 100644 index 0000000..302d678 --- /dev/null +++ b/sast-files/.agents/skills/sast-fileupload/SKILL.md @@ -0,0 +1,557 @@ +--- +name: sast-fileupload +description: >- + Detect insecure file upload vulnerabilities in a codebase using a two-phase + approach: first find all file upload handling sites (endpoints, storage calls, + multipart form processing), then check whether an attacker can upload malicious + files by manipulating file extensions. Requires sast/architecture.md (run + sast-analysis first). Outputs findings to sast/fileupload-results.md. Use when + asked to find file upload, unrestricted upload, or extension bypass bugs. +--- + +# Insecure File Upload Detection + +You are performing a focused security assessment to find insecure file upload vulnerabilities in a codebase. This skill uses a two-phase approach with subagents: **discovery** (find all places where uploaded files are received and stored) then **bypass** (determine whether an attacker can upload a malicious file by manipulating its extension or bypassing validation logic). + +**Prerequisites**: `sast/architecture.md` must exist. Run the analysis skill first if it doesn't. + +--- + +## What is an Insecure File Upload + +Insecure file upload occurs when an application accepts files from users without properly validating or restricting what can be uploaded, allowing an attacker to upload executable or malicious files. The most critical outcome is **Remote Code Execution (RCE)**: an attacker uploads a web shell (e.g., a `.php` file) and the server executes it when accessed via a direct URL. + +The core pattern: *a user-supplied file reaches a storage location without adequate extension validation, and the stored file is accessible or executable.* + +### What Insecure File Upload IS + +- Accepting any file type with no extension or content check: `file.save(upload_path)` with no validation +- Content-Type-only validation: checking `Content-Type: image/png` without verifying the actual extension or file content — trivially bypassed by setting the header manually +- Extension blocklist with gaps: `.php` is blocked but `.php3`, `.php4`, `.php5`, `.phtml`, `.phar`, `.shtml` are not +- Case-insensitive bypass: blocking `.php` but allowing `.PHP`, `.Php`, `.pHp` +- Double extension bypass: `shell.php.jpg` — code extracts the last `.jpg` and considers it safe, but the server (Apache) serves it as PHP +- Path traversal in filenames: `../../webroot/shell.php` stored via an unsanitized filename +- Incomplete filename sanitization: only stripping `../` but not encoded variants `%2e%2e%2f` +- Serving uploaded files from a web-executable directory without disabling execution + +### What Insecure File Upload is NOT + +Do not flag these as file upload vulnerabilities: + +- **Stored XSS via SVG**: uploading an SVG with embedded `" +> Or: Visit https://app.example.com/# and observe alert box] +> ``` +> +> ### [LIKELY VULNERABLE] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **XSS type**: [Reflected / Stored / DOM-based] +> - **Issue**: [e.g., "Stored user bio likely rendered via innerHTML; write path confirmed from user input"] +> - **Taint trace**: [Best-effort trace, with uncertain steps identified] +> - **Concern**: [Why it's still a risk — e.g., "Sanitization library present but configured to allow script-capable tags"] +> - **Remediation**: [Specific fix] +> - **Dynamic Test**: +> ``` +> [payload to attempt] +> ``` +> +> ### [NOT VULNERABLE] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **Reason**: [e.g., "Output wrapped in htmlspecialchars() before echo" or "Variable is a hardcoded server constant"] +> +> ### [NEEDS MANUAL REVIEW] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **Uncertainty**: [Why the variable's origin or escaping status could not be determined] +> - **Suggestion**: [What to trace manually — e.g., "Follow `buildProfileHtml()` in utils.js to check where its return value originates"] +> ``` + +--- + +## Important Reminders + +- Read `sast/architecture.md` and pass its content to both subagents as context. +- Phase 2 must run AFTER Phase 1 completes — it depends on the recon output. +- **Phase 1 is purely structural**: flag any dynamic variable passed to an HTML/JS/DOM sink, regardless of origin. Do not attempt to trace user input in Phase 1 — that is Phase 2's job. +- **Phase 2 is purely taint analysis**: for each sink found in Phase 1, trace the variable back to its origin. If it comes from a user-controlled source with no effective escaping, the site is a real vulnerability. +- Context matters: the same variable may be safe in one output context (HTML body with escaping) and dangerous in another (JavaScript string literal, URL attribute, or event handler attribute). Check the exact rendering context. +- Custom sanitization (homegrown regex stripping, blacklisting `" +> Or: Visit https://app.example.com/# and observe alert box] +> ``` +> +> ### [LIKELY VULNERABLE] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **XSS type**: [Reflected / Stored / DOM-based] +> - **Issue**: [e.g., "Stored user bio likely rendered via innerHTML; write path confirmed from user input"] +> - **Taint trace**: [Best-effort trace, with uncertain steps identified] +> - **Concern**: [Why it's still a risk — e.g., "Sanitization library present but configured to allow script-capable tags"] +> - **Remediation**: [Specific fix] +> - **Dynamic Test**: +> ``` +> [payload to attempt] +> ``` +> +> ### [NOT VULNERABLE] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **Reason**: [e.g., "Output wrapped in htmlspecialchars() before echo" or "Variable is a hardcoded server constant"] +> +> ### [NEEDS MANUAL REVIEW] Descriptive name +> - **File**: `path/to/file.ext` (lines X-Y) +> - **Endpoint / function / component**: [route, function, or component name] +> - **Uncertainty**: [Why the variable's origin or escaping status could not be determined] +> - **Suggestion**: [What to trace manually — e.g., "Follow `buildProfileHtml()` in utils.js to check where its return value originates"] +> ``` + +--- + +## Important Reminders + +- Read `sast/architecture.md` and pass its content to both subagents as context. +- Phase 2 must run AFTER Phase 1 completes — it depends on the recon output. +- **Phase 1 is purely structural**: flag any dynamic variable passed to an HTML/JS/DOM sink, regardless of origin. Do not attempt to trace user input in Phase 1 — that is Phase 2's job. +- **Phase 2 is purely taint analysis**: for each sink found in Phase 1, trace the variable back to its origin. If it comes from a user-controlled source with no effective escaping, the site is a real vulnerability. +- Context matters: the same variable may be safe in one output context (HTML body with escaping) and dangerous in another (JavaScript string literal, URL attribute, or event handler attribute). Check the exact rendering context. +- Custom sanitization (homegrown regex stripping, blacklisting `