You are probably here for one of three reasons. You saw “Sounder-Healthcare-Privacy-Scanner” in your web server logs and want to know what it was doing. You got one of our scan reports and want to know how we arrived at it. Or you are a reporter or researcher checking our work before you rely on it.

All three get the same answer. This page explains what the scanner does, what it deliberately does not do, and where it can be wrong.

Who we are

Sounder is a Chicago company that helps healthcare organizations find and remove tracking technology that shares visitor data with third parties. We are part of Pilot Digital, a full-service digital marketing firm with a large healthcare vertical. We built the scanner for our own consulting work and then made it public, because the fastest way to explain the problem is to show a clinic or hospital the issues with its own website.

Three kinds of scans

The scanner runs in three situations, and it behaves a little differently in each.

Public scans. Anyone can request a free report on a website at sounderdata.com. We assume the person asking has a connection to the site (an owner, a marketer, an agency, a compliance officer). We scan any one site at most twice per hour so the tool can’t be used to hammer a website.

Monitoring scans. Clients who hire us get their site scanned weekly, and we tell them when something changes. A new tag that shows up on a Tuesday should not wait for the next annual audit.

Research scans. We are studying how well clinics and hospitals follow privacy law, and that means scanning sites nobody asked us to scan. For those, the scanner announces itself. Every request carries this User-Agent string:

Sounder-Healthcare-Privacy-Scanner/1.0 (+https://sounderdata.com/research/methodology; research@sounderdata.com)

That string is likely how you found this page. Research scans also read and obey all robots.txt files. If it tells us to stay out, we stay out, and the site is recorded as “declined” rather than scanned.

Public and monitoring scans, where the site owner asked for the scan, use an ordinary Chrome browser identity instead. The point of those scans is to see exactly what a visitor sees, and some tracking tools behave differently when they know a robot is looking.

What a scan actually does

The scanner is a real, current Google Chrome browser controlled by software. It is not a script that downloads HTML and searches it for keywords. That matters because most tracking today is added by JavaScript after the page loads, and a keyword search never sees it.

Here is the sequence for one site.

  1. Load the homepage and wait for it to finish, including the scripts that run after the page appears. We give each page up to 30 seconds.
  2. Record everything that happened on its own. Which third-party servers did the page contact? Which cookies were set? What was written to the browser’s local storage? We take this snapshot before anyone has touched a cookie banner.
  3. Look for a consent banner and click “Accept.” We accept everything, the way most visitors do, using the banner’s own buttons. Then we record again. Anything new in the second recording is marked as having fired only after acceptance.
  4. Follow links to other pages on the same site, up to 50 pages for a standard report. Contact, appointment, and patient portal pages get priority because they are where visitors type in sensitive things. The cookie banner is accepted once per scan, so these interior pages are visited in the “accepted” state.
  5. Match what we saw against a list of known tracking services. We currently check 18 categories: Google Analytics, Google Ads, Meta (Facebook) Pixel, LinkedIn Insight Tag, TikTok Pixel, Twitter/X Pixel, Microsoft Advertising, Pinterest Tag, HubSpot Analytics, product analytics tools like Mixpanel and Amplitude, heatmap and session recording tools, live chat widgets, online scheduling tools, embedded YouTube videos, Google Maps, third-party fonts, browser storage tracking, and web forms.
  6. Note where forms and trackers share a page. A form on a page with an active advertising pixel is the combination that has produced most of the lawsuits, so we call it out separately.

The whole thing takes one to five minutes per site.

What the scanner does not do

This list is short on purpose, because the things on it are the things people worry about.

  • It never fills in or submits a form. It notes that a form exists and attempts to determine which platform built it.
  • It never logs in to anything. It only sees what an anonymous visitor sees.
  • It does not collect names, email addresses, medical information, or anything else about a website’s visitors. It never sees visitors at all. It records what a page does when the scanner itself is the visitor.
  • It does not save copies of your pages. It keeps the list of third parties contacted, the cookie names and domains, and the tracker categories found. Page text is analyzed and discarded.
  • It does not attempt to get around anything. If your site blocks automated browsers, the scan fails and the report says so. We don’t use proxies, rotating addresses, or captcha services, and we won’t.
  • It does not scan a site more than twice an hour, no matter how many people ask.

How to recognize us in your logs

Research scans always send the User-Agent string above. All scans come from a single fixed address, so you can allow it, block it, or watch for it. If you block it, the scan simply fails. We would rather you write to us first, but that is up to you.

What the report means

A report lists which of the 18 categories were found, on how many pages, and how risky each one is. Two labels appear under the tracker names.

Fired before consent means the tracker loaded on the homepage before anyone touched the cookie banner. Fired only after Accept means it appeared only after the banner was accepted. Both are findings. Under HIPAA, the FTC Act, and most state privacy laws, a visitor clicking “Accept” does not (and cannot) authorize a healthcare website to send that visitor’s data to a company that has not signed a Business Associate Agreement. The distinction matters mainly under wiretap, eavesdropping, and common law statutes, where whether the visitor agreed can be part of the analysis. We show both so you and your lawyer can decide what matters for you.

The letter grade starts at 100 points. Each high-risk tracker found costs 10 points, each low-risk one 3. Google Analytics, Meta Pixel, or Google Ads firing before consent is an automatic F, because the visitor never got a say and each of those services sends identifiable data to a company that will not sign a BAA.

Where the scanner can be wrong

We have put a lot of work into this scanner and tried to make it accurate, but it can make mistakes and it has limitations:

  • The before-consent test only runs on the homepage. A tracker that appears only on interior pages will be labeled “after Accept” because we had already accepted the banner by the time we got there. It may well fire before consent for someone who lands on that page directly.
  • Timing. Some trackers fire only when a visitor scrolls, waits, or moves the mouse. We wait, but we don’t scroll or wander. Those can be missed. A miss is a false negative. We do not report trackers we did not see.
  • Blocked or slow sites. When a site’s bot protection stops us, or a page takes longer than 30 seconds to load, the report says the page could not be analyzed. It does not say the page was clean. If a report shows very few pages scanned, that is the reason.
  • Categories are broad. “Web Forms” flags a form’s existence, not proof that it collects health information. A newsletter signup and an appointment request look the same to a scanner. Read that row with your own site in mind.
  • Consent banners we can’t operate. If a banner uses a mechanism our scanner doesn’t recognize, everything is recorded as before consent, and the report tells you no banner was accepted.

If you believe a report is wrong about your site, reply to the report email. We will look at it by hand.

The research project

We are building a public picture of how healthcare websites handle visitor privacy. The scanner records, for every site it visits, which trackers were present, whether they waited for consent, and whether the site let us in at all. We publish that last number on purpose. Most web measurement studies never say how many sites they failed to reach, and we think that figure is part of the result.

Results will be reported in aggregate. We do not publish individual site findings from research scans without contacting the organization first.

The methodology on this page is versioned. Reports record the version in force when the scan ran, so a scan from one year can be compared honestly with a scan from another. This is version 0.1. Changes will be listed here.

Questions or opting out

Write to hello@sounderdata.com and tell us the domain. We will exclude it from research scans and confirm by email. Adding a rule for our User-Agent to your robots.txt works too and takes effect on the next scan.

This page describes a software tool. It is not legal advice. Whether a particular tracker on a particular page creates legal exposure depends on facts we cannot see from outside, and on the laws that apply to you. Talk to your counsel.