Hacker Newsnew | past | comments | ask | show | jobs | submit | smartdev01's commentslogin

Your PDF analysis report post caught my attention. I want to understand your research methodology of this study.

Dual Lab analyzed the complete June 2026 Common Crawl dataset (CC-MAIN-2026-25), containing 20,578,394 PDF documents. Please read more https://pdf4wcag.com/blog-news/PDF-trends-2026Q2-by-dual-lab...

Important topic

Interesting!


Please, select your preferred language in the upper right corner https://pdf4wcag.com/


Dual Lab announces the release of PDF4WCAG Accessibility Checker 1.12. The release introduces improvements to accessibility and keyboard navigation, error visualization, validation accuracy, localization to Korean language and application stability.


PDF file sizes follow an approximately log-normal distribution, making logarithmic visualization and median-based statistics more appropriate than arithmetic averages. Median file size increased steadily until approximately 2021, after which growth flattened and became slightly negative. Quadratic regression confirms that this slowdown is highly statistically significant. The evidence suggests that the typical PDF published on the web has not become substantially larger over time. Despite initial guesses, Tagged PDFs turn out to be smaller in average than Untagged ones. https://pdf4wcag.com/blog-news/analysis-pdf-file-size


All Report are available https://pdf4wcag.com/blog-news/


We only determined the technical characteristics of the documents.


This first part of the June 2026 Common Crawl PDFs analysis reveals several long-term characteristics of PDF usage on the public web:

Most PDFs remain relatively small, short documents. Encryption is uncommon and generally does not prevent document access. Accessibility text extraction is enabled in most encrypted documents, although a significant minority still disables it. PDF 1.7 continues to dominate document production. Proprietary extensions remain common, particularly in annotation workflows. Link annotations dominate all other annotation types combined.


The post says:

> Because Common Crawl stores only the first 1 MB of each PDF

That limit became 5 MB in March 2025.


For each PDF we extracted:

basic metadata: page count file size creation and modification dates PDF version (including Version entry in the document catalog) producer and creator encryption information and permissions annotations presence of interactive forms presence of optional content layers presence of digital signatures image only (scanned) pages Tagged PDF information: stats on the use of structure element types logical structure tree validation against ISO 32005


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: