Identifying Buying Intent Signals From Public Web Activity
How to spot buying signals buried in public web behavior before deals surface.

A pricing page visit doesn't mean anything by itself. Nobody at a target account sends up a flare that says "we're buying now." Buying intent gets pieced together from scraps: a job posting here, a funding round there, three people from the same company reading the same case study in one week. This piece maps those scraps, the tiers they fall into, and the pipeline that turns them into something a sales team can act on.
Why buying intent is assembled, not declared
Timing is the whole game in B2B sales, and most teams lose it before they even know the deal exists. Buyers now finish most of their research (comparing vendors, reading reviews, asking peers) before a sales rep ever hears from them. By the time someone fills out a form, the decision is close to made. That leaves an earlier window where the account is thinking hard about a problem but hasn't told anyone yet, and that window is where deals actually get won or lost.
An intent signal is a behavioral clue that someone at a target account is actively wrestling with a problem a given product solves. No single clue proves that. A visit to a pricing page is mildly interesting on its own. That same visit paired with a job posting for a relevant role, and a funding round that just landed, turns mildly interesting into "call this account today." One data point is a guess. A combination is a pattern, and patterns are what the earliest-moving teams learn to read before a prospect ever surfaces on a competitor's radar.
How the public web shrinks the visible signal surface
The signal surface itself is shifting, and it's shrinking, not vanishing. A growing share of decision-makers now use AI tools to research a category, compare vendors, and get a summary without clicking through to a single vendor site. The page views that used to land in a marketing team's analytics dashboard are increasingly getting swallowed by an AI assistant that answers the question before the buyer ever leaves the chat window.
The scale of that shift is measurable. Per Adobe Analytics data cited in the research brief, AI-referred traffic to retail sites rose sharply year over year into 2026: AI agents are now a first-class traffic source, not a rounding error. For a sales or marketing team watching only first-party website analytics, that means a shrinking slice of the actual research happening out there.
None of this makes intent detection impossible. It changes the mix. As first-party signals get thinner, third-party intent data and public web signals, the kind that can be collected without the buyer ever landing on a company's own site, carry proportionally more of the weight. That's the case for building signal collection now instead of waiting for the picture to get clearer, because it won't.
The three tiers of public intent signals and what each tier measures
Public intent signals split into three tiers, and each one trades fidelity for coverage differently. Mixing them up, treating a tier-three signal like it carries the same weight as a tier-one one, is where most disappointment with intent data actually comes from.
Tier one covers first-party behavioral signals: pricing page visits, repeat product page views, return visits from the same account inside a short window, case study reads, ROI calculator use, content downloads. OneAway's 2026 analysis calls the pricing page visit the highest-intent signal a company can get without an actual form fill. These signals carry the most weight because they're specific to a vendor's own funnel, but they only cover accounts that already know the vendor exists, which is a real ceiling on how much of the market they can ever show. One useful pattern inside this tier: when three or more people from the same account show up within a week, that's internal sharing happening, which reads as a buying committee forming rather than one curious employee.
Tier two is second-party data, signals living on platforms a vendor doesn't own but can watch: G2 category page activity, LinkedIn engagement, review site comparisons. These sit between pure education and active evaluation, which is a meaningfully different stage than a first-time pricing page visit. The value here is reach. Tier two picks up accounts that are actively comparing options but haven't touched the vendor's own site yet, extending visibility past what first-party tracking alone can ever show.
Tier three is the widest net: aggregated co-op data, primarily from Bombora and licensed out through platforms like 6sense and Demandbase, tracking content consumption across thousands of B2B publications and ranking it by topic, alongside real-time public web signals that anyone willing to build a collection system can pull directly: job postings, funding news, technographic shifts, community chatter. This tier trades fidelity for scale. It's the best tool for finding accounts that haven't engaged with anything a vendor owns or partners with yet, even if any single signal inside it is weaker on its own.
The specific public web signals that reveal account-level buying behavior
The signals worth collecting from the open web show organizational change, real budget commitment, or active category research, not just an employee reading an article on their lunch break.
Job postings are one of the clearest. A job listing is, functionally, a public announcement of where budget is going. Role titles, the specific language in a job description, and how fast a company is hiring for a function together point at what problem that company is trying to solve, and by extension what it might be shopping for. Pairing real-time tech stack detection with hiring velocity from applicant tracking systems can produce an actual account-level score rather than a guess. That said, this data has real limits: a test run on the research brief's dataset, sampled from an August 2026 Hacker News thread and covering 111 companies, showed strong but imperfect field-level parsing. Useful, not flawless.
Funding and leadership changes carry similar weight. A funding round means budget just became available and there's now a mandate to grow. A new VP hire means there's a new decision-maker with no existing vendor loyalty. OneAway's 2026 analysis flags one compound version of this worth acting on fast: a new VP of Sales at a target account who then follows the vendor's CEO on LinkedIn. That combination decays within weeks as the account moves through its own evaluation process, so acting on it fast matters more here than almost anywhere else in the stack.
Review site behavior tells its own story. Someone reading G2 category pages is shortlisting vendors, not casually browsing. Someone reading a competitor's reviews is running switching-cost math in their head.
Technographic shifts round it out: a company swapping out one tool for another often triggers evaluation of everything adjacent to it, since a new CRM tends to drag a wave of integration and workflow tool decisions behind it. Picking this up requires actually scraping or pulling from public-facing technology indicators: script tags on a site, tech mentioned inside job listings, references buried in product documentation. Add in social engagement, likes and comments on posts about a competitor's shortcomings or a switching story: it points to someone deep in category research and willing to say so in public.
Signal stacks versus individual signals as intelligence
Take that G2 visit from a few paragraphs ago. On its own, it's a shrug. Add a job posting for a role that implies fresh budget, plus a Series B that closed last month, and the account moves from a shrug to a call worth making today. Each element doesn't just add to the others, it multiplies them, because together they answer three separate questions at once: is this company looking, can they pay, and is there a reason for them to move now.
The mechanism described in Jeeva AI's 2026 analysis evaluates combinations, frequency, and recency together rather than any one action in isolation. Recency carries its own weight separate from the combination itself. OneAway's analysis found that a content download from 18 months ago carries little weight as a signal. That same download from yesterday, followed by a case study view today, tells them a lot. Frequency matters on its own too: multiple people from one account hitting the same content in a short window reads as internal sharing, and a group is forming around this decision rather than one person doing homework alone.
Scoring models built on historical deal data assign weight to these clusters and keep adjusting as new signals roll in, so the model adjusts continuously rather than following a fixed rulebook. A company can trip several signals without being a real buyer at all: a competitor doing recon, a grad student writing a thesis, or a journalist background-checking a story can all look like a hot account on paper. The fix is filtering every signal stack through ICP fit first. A stack of signals only means something once the account actually matches the profile that has historically turned into revenue.
The structural flaw in how most teams use intent data once they have it
Good signals get wasted constantly because of what happens after collection. It's in what happens after: teams pull a weekly list of high-intent accounts and drop every one of them into the same generic outreach sequence a cold list would get. At OneAway's 2026 benchmarks, that approach produces a 0.8% meeting-booked rate, which is close to what a blind cold-email sequence gets anyway. All that signal collection, and the message on the other end never even mentions what triggered it.
If the email a high-intent account receives reads identically to what a cold account gets, the signal did one job (getting the account onto a list) and skipped the job that actually mattered (shaping what gets said to them).
A second failure occurs in vendor selection for the data itself: leaning too hard on third-party co-op signals, the Bombora-style aggregated feeds, without checking them against first-party behavior, produces lists full of accounts that look hot on paper and never close. And there's a coverage gap no amount of clever scraping fixes: buying committees increasingly hash things out in private Slack communities, Discord servers, and LinkedIn DMs, spaces that produce zero scrapable signal. Public web signals will always show a partial picture of the buying process, never the whole thing.
The fix is straightforward, even if it takes discipline to execute. The specific signal that triggered outreach should shape the actual message: the opening line, the case study referenced, the pain point named out loud. OneAway's own example makes the case: a rep referenced the exact case study a prospect had viewed, and closed a deal within three weeks. That's the entire value of intent data, realized in one sentence of a cold email that wasn't actually cold.
Building a pipeline that collects, scores, and acts on public web signals at scale
A production-grade intent pipeline runs on four layers: collection, enrichment, scoring, and activation. Skipping deliberate design on any one layer reintroduces the execution failures from the last section, just with better data feeding them.
Collection starts with two separate jobs that get confused constantly. A search layer handles discovery, finding the URLs worth looking at: job listings, funding announcements, review pages. A separate scraping layer then reads the actual content of those pages once discovery has pointed at them. For anything feeding an LLM or a retrieval pipeline, the output needs to be clean Markdown or structured JSON instead of raw HTML cluttered with navigation bars and scripts. The dual-format practitioner consensus is Markdown for chunking and embedding in retrieval pipelines, and JSON for populating the structured records the application queries directly.
Enrichment ties a raw job posting URL or funding announcement to an actual account record, decorated with firmographic and technographic context so it means something to a scoring model. The pipeline typically hands off to tools like Clay, HubSpot, Salesforce, or Zapier, which take the enriched, structured record and make it usable downstream.
Scoring should combine signal type, recency, frequency, and ICP fit into one weighted number instead of a raw count of how many actions an account took. An intentScore of 0-100 per account, updated continuously as new signals arrive, is the production pattern described in the research brief. LLMs can help at this layer, but only when they're working off retrieved, current data fed to them directly, since they distort or invent details on recent or niche topics when left to rely on their own training memory.
Activation closes the loop: the specific signals that raised an account's score should shape what the outreach actually says. Picture a mid-market account that posts a senior role signaling adjacent buying behavior, gets flagged by a technographic scan for adding a new CRM the same month, and shows two people from its domain reading the same competitor comparison page within four days. Public intent signals divide into three tiers with meaningfully different fidelity, latency, and coverage profiles, and conflating them is the source of most intent data disappointment. Together, run through a scoring layer that weights recency and frequency, they produce an account worth prioritizing today, with a message that can reference the CRM switch directly instead of guessing at a pain point.
Anti-bot infrastructure as the constraint on running the pipeline at scale
None of the four layers matter if the collection layer can't actually reach the pages it needs. Job boards, review sites, and social platforms all have a financial incentive to block automated traffic, since scraping consumes their infrastructure without paying for it the way a human visitor's ad impressions do.
This is an infrastructure problem before it's a data problem. A pipeline that scores accounts perfectly but only manages to actually retrieve half its target pages on a given day produces incomplete, unreliable intent scores, which then feed unreliable outreach decisions. Anti-bot defenses have only gotten more sophisticated, and the sites that carry the richest signals, job boards, review platforms, funding databases, tend to be exactly the ones investing the most in blocking automated access. Building a pipeline that survives contact with that reality is a separate, ongoing engineering problem from the scoring logic itself, and it's the one that decides whether everything described above runs reliably at scale or breaks quietly in ways a sales team won't notice until the pipeline of accounts starts drying up.


