Open source · MIT · command line

breaklint checks the page that will actually be printed.

When you generate PDFs from HTML, the mistakes only appear after pagination: a heading stranded at the foot of a page, one line of a paragraph carried alone onto the next, a block that fits on no page at all. breaklint measures every page after the break and reports, for each finding, the measured value next to the threshold it failed.

Source on GitHub Package on npm

Film · 0:39 · Instrumental music, no voice-over

What the film shows
  1. 0:00 Perfect in the browser. Broken on page 47. · Heading alone at the bottom · Page 48: two lines. That's it. · Table fits on no page · Client: “What happened on page 47?” · Fri 17:58 · off to the printer · Chart label cut off · Did anyone check the PDF? · It looked fine in the browser. · Page: 1 → 212 / 212
  2. 0:06 212 pages. Check them by hand? Every single build?
  3. 0:10 One command. Every page measured.
  4. 0:13 Open source · MIT · breaklint · Layout checks for HTML to PDF, measured after pagination.
  5. 0:17 It tells you where it breaks. · Real output
  6. 0:21 Page 1: heading stranded. · Sample · page 1
  7. 0:25 Page 4: 5 % filled. · Sample · page 4
  8. 0:29 Every finding shows its numbers. · Widows & orphans · Stranded headings · Blocks too tall to fit · SVG text overflow
  9. 0:33 Catch it before your client does. · Free · open source · MIT · npx breaklint --demo

Try it in one line

npx breaklint --demo

That command needs no browser and no configuration. It ends with exit 1, because the bundled fixture contains findings on purpose — a demo that ends 0 never shows you what a finding looks like. The block below shows one complete finding from the bundled demo; only the long lines are wrapped to fit the reading width. The rule chain running in it is the real one; the page it judges is a hand-written snapshot, built so that every rule path is reachable without a browser — and the report states which kind of fixture it read in its own source field.

error svg/text-overflows-viewport  page 5
  measured   72 px; threshold 0 px (uncalibrated)
  detail     This text extends 72.00 px beyond the
             SVG viewport and is not drawn.
             Coordinates are normalised through
             getScreenCTM().
  remedy     Text rendered inside an SVG extends
             outside the SVG viewport bounds and is
             clipped. Enlarge the SVG 'viewBox' or
             its width/height, or adjust the <text>
             coordinates ('x', 'y', 'text-anchor').
             'overflow: visible' on the container
             also clears the finding, but it does not
             move the text: the viewport then no
             longer clips, the target becomes
             non-applicable and this rule stops
             measuring it. Use that only where the
             overflow is intended.
             untested: no trigger/remedied pair in
             this package shows this advice removing
             this finding
  source     unknown (node produced by the paginator)
  render     unknown (no evidence produced)

Why this exists

Prose linters never see a page. Typesetting systems see the page but do not know German typographic convention. For documents generated from HTML to PDF, neither exists.

The reason for measuring the rendered page rather than the source is a case that source-level checking did not catch: a chart passed an XML validity check and a geometry check, and still came out of the renderer with colliding labels. Whether a page works is decided after pagination — in the renderer, with the fonts that were actually available.

What is checked

Fifteen rules, twelve of them active by default. Two error rules can fail a build by default; the ten other active rules are advisory until you ask for more with --fail-on warn. The two new figure rules are warnings (warn), default-off and uncalibrated. layout/half-empty-page remains experimental, never moves an exit code and has been outside the default profile since 0.6.0. These three rules can run with --profile strict or explicit enablement.

Pagination

Unbreakable blocks taller than the page, widows and orphans, headings at the foot of a page, orphaned continuation pages, and hyphenation across a break. Half-empty pages only when you ask for them.

SVG

Text that runs past the viewport. Clip/mask removal and device-pixel collision are not released rules yet: production has no safe ink pass for them.

Typography

A hyphen where a dash belongs, typewriter quotes in typeset prose, short last lines, and word gaps torn open by justification.

Artefacts

file: URIs and build-machine paths left behind in the delivered document.

Figure, caption and local reference

figure/caption-separated measures the page distance between an HTML figure's image body and its visible figcaption. It supports one img or single inline SVG with exactly one owned caption. Split or repeated bodies, multiple captions and missing geometry are not measured. A finding proves neither the pagination cause nor that the figure and caption will fit on one page.

figure/dangling-reference checks local href anchors in links whose text contains “Figure”, “Fig.”, “Abbildung” or “Abb.” followed by a number. It reports absent target IDs and identifies the containing source block, not the exact inline link position. External links, free-text numbering and an existing target's visibility are not validated. This is not a complete figure model.

Working through findings

Human reports show up to three next checks. Partial-run measurement limits come before findings, followed by checks of the available results. Counted decline reasons remain visible even when the coverage floor is fulfilled. The HTML rule overview provides navigation; individual findings remain in the report.

JSON remains canonical. Console, HTML and Markdown derive their guidance from it without inventing a cause or reader impact. Markdown includes the actual finding message and remediation advice with its tested flag. Untested advice is not evidence that a change works. Baseline comparison remains bound to verified, compatible measurements.

What remains uncalibrated

All fifteen released rules are still labelled uncalibrated. The two ink definitions remain explicitly unreleased M3 research. Evidence for later calibration includes runs in the real renderer and a rights- and privacy-reviewed process pilot with three real documents: separate data for development, threshold tuning and the final test, plus blind review packets and controls that reject invalid evidence.

That is not empirical calibration yet. It still requires a sufficiently large corpus of real documents, two independent blind human judgements, a documented resolution when they disagree, an externally verified final test set that has remained unchanged, and one final evaluation defined in advance. Until then, every rule remains calibrated: false.

Tests and renderer evidence check documented measurement paths and known counterexamples. A mandatory real-document gate binds an unchanged, rights- and privacy-reviewed third-party Project Gutenberg HTML file plus a first-party packaging case, requires exact page and measured-rule counts, and rejects infrastructure drift; that proves robustness, not calibration. How often heuristic boundaries are correct across real documents remains unknown.

What it does not do

No PDF standard conformity

Use veraPDF or pdfcpu for that.

No visual image comparison

Use a visual-diff tool for pixel or image differences. breaklint's conservative document-repair comparison instead requires a complete, verified compatible target.

No prose or style checking

Use vale or typopo for that.

PDF is not an input format

PDF is produced and rasterised here to make evidence, never read as input.

Requirements

Current release: 0.9.0; 0.8.0 is its immediate predecessor. Fifteen rules are registered, twelve active by default. All rules remain uncalibrated. A failed image becomes the non-fatal image-content-unavailable diagnostic only when positive authored width and height attributes exactly equal the measured browser box; replacement text and CSS drift are insufficient.

Since 0.8.0, each document has a default time budget of 600000 ms (ten minutes). Set documentTimeoutMs in breaklint.config.json, or override it for a run with --document-timeout-ms <n>. The value must be an integer from 30000 to 1800000 ms. Invalid values exit 2. A run that reaches the budget exits 3 and names the effective value. The budget limits duration; it does not speed up pagination or reduce the memory use of large documents.

0.8.0 refuses two known paths to silent content loss before measurement: a body with computed column-count or column-width (including column-count: 1), and a nonempty heading with display: contents. Both exit 3 instead of reporting clean. Repeated SVG text in ordinary page content is assigned by occurrence. Inline SVG text in running margin boxes remains an exception: when page copies share a target ID across pages, the check ends in checker-crashed (exit 3). This does not establish a general PDF content-completeness check.

For a narrow class of visible strokes around SVG text, breaklint proves that a conservative envelope lies strictly inside the SVG viewport. This also requires supported stroke properties and transforms. If the envelope touches the viewport edge, the target remains unmeasured and svg/text-overflows-viewport exits 4 because it requires full coverage. This proves containment, not exact painted bounds or an overflow finding. Masks and filters remain unmeasured.

0.9.0 writes Snapshot 6 with the bounded figure index. Report 5, agent context 2 and Configuration 1 remain unchanged. Stored version 5 snapshots remain readable; when figure rules are enabled, their missing index causes a measurement decline with env/figure-index-unavailable.

Since 0.7.0, layout/unbreakable-block-too-tall, one of the two rules that fail a build by default, sums a block's height over its fragments once the paginator has split it into three or more pieces; a too-tall block split into exactly two pieces is still not reported. A split block whose fragments cannot be joined by their source is not measured at its first fragment; the rule declines it and counts that against coverage, so such a run can end with exit 4. Content in the page margins, such as running headers and footers from position: running(...), does not count as text flow and is judged by no rule. Whether a page boundary was forced is read from the paginator's own break decision, including in regions with named pages. A report written to a pipe arrives whole; through 0.6.0 it could be cut at a multiple of the pipe buffer, on Linux usually at 64 KiB. If it cannot be written completely, the run ends with exit 3.

For produced documents, breaklint names an editable original location only when a host-controlled producer has captured the original bytes and proved the source relation. The same canonical JSON produces an offline human view and bounded context for AI systems. A repair comparison marks a finding resolved only when the target is verified and identical, and the revision, rules, configuration, renderer environment, fonts and resources are compatible; unknown sources and partial coverage remain explicit.

The GitHub Action introduced in 0.7.0 can be included as uses: godarg/breaklint@v0.9.0. It installs the published npm release, checks the given HTML paths in one run, writes SARIF, JUnit and a Markdown summary from the one canonical JSON report, and ends the step with breaklint's exit code; exits 2, 3 and 4 always fail it. docs/ci-recipe.md has a workflow to copy. The Action is tested on ubuntu-latest; macOS runners are untested. It needs bash 4 or newer; with macOS's /bin/bash 3.2 it stops with exit 3.

Separately, the public checkPage API checks an already prepared Playwright Page; it does not take over navigation, authentication or network policy. In our own operations it checks 45 states in an internal dashboard's test suite. The command-line path still accepts only .html/.htm; standalone SVG, PDF and Markdown files end with usage Exit 2. Node 22.13 or newer, on macOS or Linux. A run over real documents additionally needs a Chromium-based browser and pagedjs@0.4.3; pdfjs-dist rasterises the produced PDF to bind evidence to findings. Windows is not supported — process termination here rests on POSIX process groups. The measurement chain is exercised on Linux; process termination and profile cleanup have empirical evidence on macOS only.

npm i -D breaklint
npm i -D puppeteer-core@^25.8.0 pagedjs@0.4.3 pdfjs-dist@6.2.108

breaklint executes foreign HTML, scripts included. Offline mode intercepts page requests but is not an egress control: it does not cover WebSocket, WebTransport or WebRTC connections a document opens, nor browser traffic through DNS over HTTPS or the component updater. Since 0.8.0, breaklint refuses to launch the renderer when PUPPETEER_DANGEROUS_NO_SANDBOX or PUPPETEER_TEST_EXPERIMENTAL_CHROME_FEATURES is present; it removes CHROME_EXTRA_FLAGS from Chrome's environment. Other wrapper variables have not been generally audited. Untrusted documents belong in a container or network namespace with no egress.

Still open: Documents with Paged.js footnotes exit 3 before the rules run. Individual rules decline to measure multi-column blocks. An oversized unbreakable block split into exactly two fragments still goes unreported.

breaklint uses no language model at check time. It performs no inference and contains no API client.

Open community test · voluntary · unpaid

Where do the measurement and a human eye disagree?

Those are exactly the cases we are looking for. Anyone can test public or purpose-built HTML — no application, selection or prior experience required. The short guide takes about ten minutes from installation to a structured report.

Use only material you have the right to publish. Your GitHub identity, report and reproduction are public; confidential documents, personal data and security findings do not belong in the form. Structured intake is automated, but submitted content is never executed.

This is additional product QA. Public testing is not a blind study and does not replace the empirical calibration that the rules still need.

Where it comes from

breaklint is built at Dargel Solutions and is free to use under the MIT licence, including commercially. Parts of the repository were written with the help of large language models. Rules, thresholds, sources and documented limits are available in the public repository. Rules that cite the German orthography ruleset were checked against its published text.

Generative AI is used to create, maintain and review this page. Dargel Solutions is responsible for publication.

breaklint is listed in the PeerPush directory.