Skip to content
yceffort
PostsSeriesTagsAbout🧪 Research
KO

Tweaks

theme
accent palette
film grain
minimal mode
◆ SERIES · 2 POSTS

OG Scraping Server Design Notes

What should you decide first when building a link preview server? Following the principles from the reasoning behind the runtime choice to SSRF, encoding, and cache stampedes, and running every piece of code written along the way

README

On paper, a link preview is a three-line feature. Open the URL, parse the HTML, pull out the og: tags. In production, though, those three lines break in quite a few places, and many of them are lumped together behind a single number, "the error rate is high", where they are hard to see.

This series is a set of design notes that follows, in order, the places where those three lines break. Part 1 splits failures by cause, examines where the runtime choice actually makes a difference, and goes on to raise coverage with User-Agent and encoding handling and to set up caching and target numbers. Part 2 deals with the risk of a server opening a URL on behalf of a user, and looks at how it gets broken before how to block it.

The focus is on why the design ends up the way it does rather than on usage, and every piece of code shown was run on Node.js v24.14.1. In the process I found that four of the things I had first written down plausibly were wrong, and I left them in place instead of deleting them. That contrast is also what I most wanted to say in this series.

02 ITEMS

All posts

in order
01

Building an OG Scraping Server in Node.js (1): From Runtime Choice to Error Rate and Latency

#web-scraping#backend#nodejs
The "10% error rate" of a link preview server is a single number that five different kinds of failure got mashed into. This post works out why this workload is I/O bound at that TPS, where runtime choice actually diverges across four points, and then moves on to lowering the error rate with User-Agent and encoding. Node built-in TextDecoder turns CP949 extension characters into different characters without raising an error, and a scraped og:title is not an API response but user input. It also covers cache stampedes, negative caching, and a two-million-run simulation that verifies "P95 under one second" by working backwards from the cache hit rate. The first post of a two-part design note on OG scraping servers.
2026-08-2233 min · read
02

Building an OG Scraping Server in Node.js (2): How SSRF Gets Through

#security#nodejs#backend
A feature where the server opens a URL the user handed it has the textbook conditions for SSRF written into its spec. Six ways a whitelist gets bypassed first, then five defensive principles that block them, all actually run on Node. Strip IPv4-mapped by hand and it gets through in hex notation, undici lookup hook is never called when the host is an IP literal, and URL.hostname keeps the brackets on an IPv6 literal. The final post of a two-part design note on OG scraping servers.
2026-08-2225 min · read
mailMail icongithubtwitter
yceffort
•
© 2026
•
https://yceffort.kr