Compliance Web Activity Archiving

Compliance Web Activity Archiving

Everything you publish is a statement your organization made, on a date, to the public. The product description, the pricing page, the terms of service, the privacy policy, the disclaimer at the bottom of a landing page. Each one is something you may later need to prove.

The gap at most organizations is that nobody can prove what the site said on a given date. Hosting backups exist, but a backup is a database dump, not a record. It is rewriteable, it carries no chain of custody, and restoring it to see one page means standing up an entire environment. WordPress post revisions exist, but they are editable by any administrator, they cover only post and page content, and they miss theme output, menus, widgets, plugin-generated content, and anything assembled by a page builder at render time. The Internet Archive exists, but it is incomplete, third-party controlled, and captures only what its crawler happened to reach on a schedule you do not set.

When someone asks you to produce a page as it appeared eighteen months ago, complete with the language that was live at the time, most organizations cannot answer. Compliance Web Archiving closes that gap.

What we deliver

We implement and operate continuous, tamper-evident archiving of your web properties, and we take responsibility for the part that usually breaks: making it work correctly against a real WordPress site.

The service includes:

  • Continuous automated capture on a schedule matched to how often the site changes
  • Tamper-evident records carrying cryptographic hash values, timestamps, and digital signatures, stored in write-once form
  • Interactive replay of any historical capture, so a record can be navigated the way a visitor experienced it rather than read as a flattened screenshot
  • Full-text search across the archive by date, keyword, or URL
  • Version comparison showing exactly what changed between any two captures
  • Change approval workflows producing a time-stamped audit trail of who approved what and when
  • Configurable retention schedules aligned to your policy or obligation
  • Legal hold applied to preserve records against deletion during litigation or investigation
  • Role-based access control governing who can view, search, and export
  • Export in PDF, CSV, and eDiscovery-compatible formats
  • Secure access portals letting auditors, counsel, or the public search records directly without routing every request through your team


Social media, text message, and collaboration platform archiving are available as additional scopes where communications outside the website create exposure.

When this matters

Organizations engage us when one or more of the following is true.

  • You have a retention obligation. Many regulated sectors require that communications with the public be preserved in complete, unalterable, promptly retrievable form. Where that obligation attaches, screenshots and backups do not satisfy it.
  • You face litigation or expect to. Web content is discoverable. Preserving it under legal hold, in a form with demonstrable integrity, is materially stronger than reconstructing it after the fact from whatever anyone happened to save.
  • Someone disputes what you published. Advertised pricing, promotional terms, product claims, availability, service descriptions. Disputes over what a page said on a specific day are common and are usually resolved by whoever has the better record.
  • You need to prove which policy version was in effect. Terms of service, privacy policies, consent language, and disclosure text change over time. Demonstrating which version a given user encountered on a given date is difficult without a dated, verifiable capture, and it comes up constantly in privacy and consumer claims.
  • You need evidence of site state at a point in time. Accessibility demand letters, tracking and consent disputes, and compliance reviews all turn on what the site actually presented and disclosed on a particular date, not what anyone remembers it presenting.
  • You need change accountability. Who changed the disclaimer, when, and who approved it. A time-stamped approval trail replaces reconstructing intent from a Slack thread six months later.
  • Your archive must be independent of your CMS. A record your own administrators can edit is a weak record. Separating preservation from publication is the point.

Why WordPress needs deliberate work

Archiving platforms are sold as a URL and a credit card. On a real WordPress site that assumption fails quietly, and a silent failure is worse than no archive at all, because you believe you are covered when you are not.

This is the part of the engagement we own.

  • Crawler access. Setup requires allowlisting the archiving crawler across every layer that can block it: your host’s firewall and bot mitigation, any CDN or WAF in front of the site, and any security plugin performing rate limiting or IP lockouts. A crawler that gets a 403 produces a gap that nothing in the platform will flag as a gap.
  • Crawl load and caching. Recurring full-site crawls generate real traffic. We schedule capture windows against your host’s resource limits and verify the crawler is reaching origin content rather than a stale cached copy.
  • Page builder and block output. Content assembled at render time by a page builder, by ACF flexible content layouts, or by dynamic blocks never appears in post revisions and may be missed by a naive crawl. We verify captured snapshots against live rendered output before signing off on onboarding.
  • Consent banners. A cookie consent overlay can obscure content in a capture or block scripts the page depends on. Since most organizations now run a consent management platform, this needs explicit testing. Your archived record should show the page as a consenting visitor saw it, not as a modal.
  • Gated content. Client portals, member areas, and logged-in resources require credentialed capture. We scope these deliberately and confirm whether gated content belongs inside your retention scope before archiving it.
  • Documents. PDFs are frequently the most consequential content on a site and the most often overlooked. We confirm linked documents are captured, not just the HTML pages linking to them.
  • Third-party embeds. Calculators, booking tools, iframed content, and hosted form platforms render from origins you do not control. We determine what is captured, what is not, and document the boundary in writing rather than letting anyone assume coverage.
  • Multi-site and multilingual. Each domain, subsite, and language variant is its own scope. A multilingual site can multiply effective page count substantially, which affects both completeness and cost.
  • Staging isolation. We confirm the crawler reaches production only, and that non-production environments stay excluded.
  • Robots directives. Archiving crawlers may or may not honor robots.txt. We verify behavior against your actual configuration so that pages excluded from search engines are not also excluded from your archive.
  • Ongoing verification. Sites change. A redesign, a plugin update, a new WAF rule, or a migration can break capture without any visible symptom. We monitor capture completeness as a standing service item rather than treating implementation as a one-time project.

What this replaces, and what it does not

Replaces: manual PDF screenshots taken before and after site updates; print-to-PDF archives on a shared drive; reliance on the Internet Archive; reliance on hosting backups as a record; spreadsheet-tracked change approvals.

Does not replace: your hosting backup, which serves disaster recovery and remains necessary. Nor does it replace email and communications archiving, records management for non-web content, or your internal governance policy. The archive is evidence. The policy governing it still has to exist, and we will not write it for you.

Security and vendor review

The archiving platform we deploy is operated by our archiving partner and carries:

  • SOC 2 Type 1 and Type 2 attestation
  • ISO 27001 certification of its management system
  • FedRAMP authorization for the website archiving solution, with continuous monitoring
  • SOC-compliant North American data centers; ISO 27001 certified Canadian and European data centers
  • Separate U.S. and EU archiving environments for data residency requirements

We supply the SOC 2 report, CAIQ documentation, subprocessor list, and completed security questionnaires to your vendor management team on request. Plan for this to take time. In larger organizations, vendor review routinely adds several weeks to the start date, and we would rather set that expectation now than miss a date later.

How an engagement runs

  • Discovery. We inventory your web properties, establish page and document counts, and work with whoever owns records or risk in your organization to confirm what needs preserving and for how long.
  • Implementation. Crawler allowlisting across host, CDN, and security layers. Capture configuration. Credentialed access for gated content where in scope. Retention and access policy configuration.
  • Validation. We compare captured snapshots against live rendered output across representative templates, confirm document capture, and document any content that cannot be captured, before you rely on the archive.
  • Handover and training. Your team learns to search, compare, export, and share records so that responding to a request does not require calling us first.
  • Ongoing service. Capture monitoring, scope updates as the site changes, re-validation after redesigns and migrations, and support during audits, reviews, and discovery.

What we need to know to scope

  1. Domains and subdomains in scope, and approximate page count including PDFs
  2. Required capture frequency, and whether it varies by section
  3. Retention period and what drives it
  4. Whether gated content is in scope, and who provides credentials
  5. Whether social or messaging channels are in scope
  6. Data residency requirements, if any
  7. Who owns the archive internally and responds to requests for records
  8. Whether an access portal is needed for third parties, and on what domain
  9. Your current WAF, CDN, and security plugin stack
  10. Whether an archiving vendor is already in place and whether historical records need migrating

Pricing is built per engagement from scope: number of properties, page and storage volume, capture frequency, retention length, and the depth of ongoing service required.

This document describes a service offering and is not legal advice. Preservation obligations vary by industry, activity, and jurisdiction. Confirm your organization’s specific requirements with qualified counsel.