Skip to main content

Command Palette

Search for a command to run...

TryHackMe: MD2PDF Writeup

Updated
•9 min read•View as Markdown
Y
I write detailed writeups on HackTheBox, PicoCTF and other CTF challenges. Passionate about web exploitation, Active Directory attacks and ethical hacking

Summary

MD2PDF is a web app that converts user-submitted Markdown into a downloadable PDF using a server-side rendering engine. The input is not sanitized, so raw HTML is rendered as-is - including tags like <iframe> that cause the renderer to make its own outbound HTTP requests. Since those requests originate from the server itself, they can be pointed at localhost, allowing an attacker to reach an /admin endpoint that was supposed to be restricted to internal/localhost traffic only. Chaining the HTML injection with this SSRF behavior let an external, unauthenticated user retrieve the contents of the restricted admin page - including the flag - entirely through the PDF export feature.


1. Reconnaissance

A port scan of the target was run first to see what was exposed.

rustscan -a machine-ip --ulimit 5000

This found three open ports, which were then fingerprinted with a full service/version scan:

nmap -sCV machine-ip -p 22,80,5000

Results:

Port State Service
22 open OpenSSH 8.2p1 (Ubuntu)
80 open HTTP - MD2PDF web app
5000 open HTTP - MD2PDF web app (same application, different port)

Both port 80 and port 5000 host the same MD2PDF instance. This detail turns out to matter later: the app itself refers to port 5000 as its "internal" / localhost-facing port.

Browsing to the site showed a simple single-page app: a textarea for Markdown input, and a "Convert to PDF" button.


2. Confirming HTML Injection

The app is advertised as a Markdown converter, but Markdown renderers commonly pass raw HTML through untouched unless explicitly configured not to. To test this, the following was submitted in the textarea instead of plain Markdown:

<b> test </best>

(note the intentionally mismatched closing tag - this was just to confirm the parser wasn't doing anything unusual with malformed input)

Clicking Convert to PDF produced a PDF where the word "test" appeared in bold. This confirmed that:

  • The input is not sanitized or HTML-escaped before rendering.
  • The conversion pipeline interprets raw HTML tags as HTML, not as literal text.

This is the entry point for the rest of the attack: anything that can be embedded in HTML and processed by a rendering engine can potentially be injected here.


3. Identifying the Rendering Engine's Capabilities (SSRF Test)

Because this is a Markdown-to-PDF converter, there's almost always a headless browser or HTML-rendering engine working server-side (e.g., a Chromium-based renderer or similar) that loads the HTML and prints it to PDF. These engines can load external resources - including iframes, images, and stylesheets - exactly like a normal browser would. Critically, when it does so, the request comes from the server, not from the attacker's machine.

To test this, an iframe pointing to an arbitrary IP was injected:

<b>
  <iframe src="http://attacker-ip"></iframe>
</b>

Converting this to PDF caused the server to actually attempt an outbound connection to that address - confirming the renderer fetches iframe src URLs server-side. This is the hallmark of an SSRF: the application will make network requests on the attacker's behalf, to destinations the attacker chooses.

A secondary attempt was made to read a local file directly via the file:// scheme:

<iframe src="file:///etc/passwd" width="800px" height="600px"></iframe>

This did not work - the renderer either blocks the file:// protocol or doesn't support it for iframes. This ruled out direct local file disclosure, but it didn't rule out SSRF against network services, which is the direction that paid off.


4. Enumerating the Application for Internal-Only Endpoints

With confirmed HTML injection, the next question was: is there anything on this server worth reaching that we can't reach directly?

A content/directory brute-force was run against the web root:

feroxbuster -u http://machine-ip/

Notable results:

403   GET   /admin
405   GET   /convert
200   GET   /static/...   (css/js assets)

The /admin path stood out - a 403 Forbidden suggests the route exists but access is being deliberately blocked. Visiting it directly in a browser confirmed this:

Forbidden
This page can only be seen internally (localhost:5000)

This is the key finding: /admin is gated by an IP/host-based check rather than real authentication. The app only serves that page when the incoming request appears to originate from localhost on port 5000 - i.e., from the server talking to itself.


5. Chaining HTML Injection + SSRF to Reach /admin

Since Step 3 established that the PDF-rendering engine makes its HTTP requests from the server itself, and Step 4 established that /admin trusts any request that looks like it came from localhost:5000, the two issues chain together perfectly:

If the renderer (running on the server) requests /admin on localhost:5000, the request genuinely does originate from localhost - the "internal only" check passes, because it isn't actually checking identity, just request origin.

The following payload was submitted to the converter:

<b>
  <iframe src="http://localhost:5000/admin"></iframe>
</b>

The server-side renderer loaded this HTML, resolved the iframe's src, and - because it was making the request locally - successfully retrieved the normally-forbidden /admin page. That rendered content was then embedded into the output PDF, which was downloadable by the attacker.


6. Result

Opening the generated PDF revealed the contents of the admin page, including the challenge flag:

flag{REDACTED}

7. Root Cause Analysis

Two separate weaknesses combined to produce this vulnerability:

  1. Unsanitized HTML input to a server-side renderer. The application accepts raw HTML (not just Markdown) and passes it directly into a rendering engine capable of making outbound HTTP requests (via <iframe>, and likely <img>, <link>, etc.). No tag/attribute allowlist or sandboxing was applied.

  2. Access control based on request origin instead of authentication. /admin is protected only by a check of where the request appears to come from (localhost:5000). This is not a security boundary when something on that same host (like the PDF renderer) can be convinced to make requests on an attacker's behalf. "Is this localhost?" is not equivalent to "is this an authenticated admin?"

Together, these allow an unauthenticated external user to indirectly make authenticated-looking requests to internal-only endpoints - a textbook SSRF-to-access-control-bypass chain.


8. Remediation Recommendations

  • Sanitize HTML input before rendering: strip or allowlist tags (removing iframe, img, object, embed, script, and similar resource-loading/network-capable elements) if the app is genuinely only meant to accept Markdown.
  • Network-isolate the renderer: run the PDF rendering process in a container/namespace with no route to internal services (deny-by-default egress, or an explicit allowlist of nothing).
  • Never trust "localhost" as an authorization mechanism. Any service that can be tricked into making a loopback request (renderers, image proxies, webhook handlers, URL preview features, etc.) defeats this check trivially. Protect sensitive endpoints with real authentication (tokens, session checks, mTLS) instead of request-origin checks.
  • Disable dangerous URL schemes (file://, gopher://, etc.) in the renderer regardless of the above, as defense in depth.

Timeline Summary

Step Action Outcome
1 Port scan (rustscan, nmap) Found ports 22, 80, 5000 - both 80/5000 run MD2PDF
2 Injected <b> tag into Markdown input Confirmed unsanitized HTML injection
3 Injected <iframe src="http://<external-ip>"> Confirmed SSRF - server makes outbound requests
4 feroxbuster directory scan Found /admin (403, "internal only")
5 Injected <iframe src="http://localhost:5000/admin"> Renderer fetched /admin locally, bypassing the restriction
6 Opened resulting PDF Flag disclosed

9. Attack Chain

A condensed view of how each weakness fed into the next:

Unsanitized Markdown input
        │
        ▼
Raw HTML is rendered server-side (HTML Injection confirmed)
        │
        ▼
<iframe> causes the renderer to make its OWN outbound request (SSRF confirmed)
        │
        ▼
feroxbuster finds /admin -> blocked unless request looks like it's from localhost:5000
        │
        ▼
Inject <iframe src="http://localhost:5000/admin">
        │
        ▼
Renderer (running ON the server) requests /admin FROM localhost
        │
        ▼
"Internal only" check passes - request genuinely IS local
        │
        ▼
/admin content is rendered into the output PDF
        │
        ▼
Attacker downloads PDF -> reads restricted admin page -> flag disclosed

In short: HTML Injection gave control over what the renderer does -> SSRF gave the ability to make the renderer issue requests on the attacker's behalf -> a flawed "localhost-only" access check was satisfied because the request really did come from localhost (just not from a trusted party).


10. Key Vulnerabilities and Mitigations

# Vulnerability Why It's Exploitable Mitigation
1 HTML Injection - the Markdown input field renders raw HTML unsanitized Lets an attacker control exactly what the server-side renderer loads and executes, instead of only safe Markdown-derived output Sanitize/strip input before rendering; allowlist safe Markdown-only elements and strip tags like <iframe>, <script>, <object>, <embed>, <link>, <img> with remote src if they aren't needed
2 Server-Side Request Forgery (SSRF) - the rendering engine fetches remote resources (e.g., iframe src) on the server's behalf Allows an external attacker to make the server issue network requests to arbitrary destinations, including internal-only services Run the renderer in a network-isolated sandbox/container with no route to internal services; enforce a strict egress allowlist (ideally none); disable dangerous schemes like file:// and gopher://
3 Broken access control on /admin - protected only by checking if the request appears to come from localhost:5000 Request-origin checks don't verify who is asking, only where the request appears to originate from - trivially spoofable by anything else running on the same host (like the renderer) Require real authentication (session tokens, credentials, mTLS) for sensitive endpoints instead of relying on source-IP/hostname checks
4 Lack of defense in depth - a single injection point led directly to full internal endpoint disclosure No compensating controls existed between "attacker can inject HTML" and "attacker can read restricted admin data" Layer controls: input sanitization AND network isolation AND proper authentication, so that no single flaw results in full compromise