NuevoVer cómo
Todas las novedades
Mejora5 min de lectura

A second reader for difficult websites

Sites that defeated the ordinary reader now get a second attempt with a browser-based one, so prospects with modern or protected websites still get properly researched emails.

When the ordinary reader gets almost nothing from a prospect's website, a second one now tries. It runs a real browser, so sites that build their content in the page still get read.

What changed

Reading a website normally means fetching the page and taking the text. That is fast, cheap, and works for most of the web.

It fails on a growing portion of it. Many modern sites send an almost empty page and then build the content in the browser using scripts. Fetch that page and you get a shell: a title, some navigation, and nothing that describes what the company does. Others sit behind protection that refuses anything that does not look like a real browser.

Either way the result is the same. Research completes, technically, having learned nothing, and the email gets written from a company name and a domain.

Now, when the first reader returns too little to be useful, a second attempt runs through an actual browser. It loads the page the way a person's browser would, waits for the content to appear, and reads what is actually on the screen.

Why it matters

The sites that defeat simple reading are not a random sample. They skew heavily towards companies that invested in a modern website, which correlates with the companies most worth talking to.

So the failure was concentrated in exactly the wrong place. The prospects who most deserved a well researched email were the ones most likely to get a generic one, and nothing about the process made that visible.

Fixing it changes the quality of emails to a specific and valuable slice of your list rather than improving everything slightly.

How to use it

Nothing to enable. The second attempt runs automatically when the first returns too little.

You will see it as better research on prospects whose sites previously produced nothing, and emails that reference something real about their business.

The lead record shows what was found, so you can see which prospects were rescued this way.

What it costs

The browser based reader is slower and more expensive, which is why it is a fallback rather than the default. Running it on every site would make research several times slower for no benefit on the majority that read fine the first time.

It only runs when the first attempt returns below a threshold of usable content. A site that reads properly never triggers it.

That keeps the cost proportional. The expensive path is spent only where the cheap path failed, which is a small fraction of sites and the fraction where the extra effort is worth something.

What it still cannot read

Content behind a login. If the useful description of a product sits behind a sign up, no reader gets it, and that is correct.

Sites that are genuinely empty. A one page site with a logo and a contact form contains no information whatever the reader, and the second attempt confirms the absence rather than fixing it.

Content in images. A page whose text is baked into a picture reads as a page with no text.

When both readers fail

The lead proceeds with what is known. Research failing is not a reason to hold a lead out of the campaign indefinitely, and after the retry ceiling the email is written from the company name, the domain, and whatever else the record already carries.

That email is thinner and it still goes. A thin email has some chance of a reply and a lead held forever has none.

The record shows research was incomplete, so if a campaign to a particular segment underperforms, you can see whether unreadable sites were part of the reason rather than guessing.

How much difference it makes

The clearest effect is on prospects whose site previously produced nothing usable. Those emails went from general to specific, which is the largest quality jump available in the whole pipeline.

It also reduced the number of leads that reach the research hold ceiling, because a lead whose site is now readable no longer exhausts its attempts on a page that was never going to yield anything.

The effect on the overall fleet is smaller than the effect on the affected leads, which is what you would expect from a fix targeted at a specific failure rather than a general improvement.

Why not use the browser reader for everything

Speed, mostly. Loading a page in a real browser and waiting for it to finish takes many times longer than fetching the same page as text, and research runs across every lead in every campaign.

Cost is the other half. A browser needs real machine resources per page, where a fetch needs almost none. Running every site through it would multiply the cost of research for a benefit confined to a minority of pages.

Using it as a fallback puts the expense exactly where the benefit is, which is the only version of this that makes sense at scale.

What good research produces

The useful output is not a summary of the website. It is a small set of specific, checkable facts about what the company does and who it serves, which is what an email can then refer to without inventing anything.

That is why a shell page is so damaging. It yields a company name and a navigation menu, neither of which supports a single sentence worth writing, so the email falls back to generalities however good the writer is.

Reading the real page changes what is available to write from, which is upstream of every other quality mechanism in the pipeline. No amount of rewriting turns absent information into present information.

Where it runs

The browser based reader runs on our own infrastructure rather than a third party service, on a host dedicated to it, so a slow page does not compete for resources with the rest of the product.

It respects the same rules as the ordinary reader: publicly reachable pages only, at a polite rate, and nothing behind authentication.

Availability

Live now on all plans, applied automatically during prospect research. No configuration and no separate charge.