How IndustryLens Collects, Verifies, and Scores Competitive Intelligence

Every claim in every briefing traces back to a specific URL, a scrape timestamp, and a recorded calculation — here's exactly what we collect, how it becomes a verified insight, and how each finding earns its confidence score.

If you read a finding you cannot verify, that is our bug. See Corrections.

350+
Data source endpoints
30
Automated quality checks
94%
Median confidence score
98%
Claim accuracy rate

350+ Data Sources — by Category

Every source is a public surface any reader can verify. We do not use data brokers, behavioural panels, or proprietary dashboards. Per observation we store: source URL, scrape timestamp, and verbatim extract — nothing is paraphrased before storage. See how source coverage works →

Review Platforms
12+ sources
G2CapterraTrustpilotTrustRadiusGartner Peer InsightsGetAppSoftware AdviceSourceForgeClutchProduct HuntAppSumoFuturepedia

Win/loss signals, feature sentiment, switching intent, named customer evidence.

Job Boards & Hiring Signals
8+ sources
LinkedIn JobsIndeedGlassdoorWellfound / AngelListBuiltinHimalayasRemote.coGreenhouse / Lever / Workday career pages

Headcount direction, new product bets, engineering priorities — 4–6 weeks before any press release.

News & PR
15+ sources
Google NewsPR NewswireBusiness WireGlobeNewswireTechCrunchVentureBeatSaaStrG2 Crowd blogCompany blogs via RSS

Funding rounds, partnerships, executive changes, and product announcements — with timestamps and verbatim quotes.

Social & Community
6+ sources
LinkedIn company postsReddit (vertical-relevant subreddits)X / TwitterYouTubeHacker News

Organic positioning shifts, community sentiment, launch traction, and feature complaints straight from users.

Ads Intelligence
4+ sources
Meta Ad LibraryGoogle Ads Transparency CenterLinkedIn Campaign Manager public dataTikTok Creative Center

Live campaign themes, ICP targeting, messaging pivots, and budget signals — all from public ad registries.

Pricing & Product
All sources
Live competitor pricing pagesFeature comparison pagesChangelog and release notesPublic product documentation

Direct-page scrapes with week-over-week diffing so every pricing change and new tier gets flagged, not inferred.

Analyst & Aggregator
5+ sources
G2 category reportsGartner Magic Quadrant updatesForrester mentionsIDC coverageCrunchbase / PitchBook funding data

Category narrative, analyst positioning, and funding context that shapes how buyers perceive competing vendors.

The 7 source categories above name ~55 platforms. The 350+ endpoint count reflects per-competitor scraping: each vendor is checked across every applicable source, and aggregator APIs (job board feeds, news RSS, ad library endpoints) count per endpoint per pipeline run — not per platform.

The 30-Check Quality Pipeline

Every observation passes through four layers of automated validation before it can become an insight in a customer briefing. Checks that fail do not gate the insight — they kill it. Nothing is downgraded silently.

A
Layer ASource Validation
5 checks
1
Source freshnessReject any observation more than 7 days old relative to the report window.
2
Source authority scoreEach source type carries a fixed credibility weight; primary-source URLs score highest.
3
Duplicate detectionHash-based deduplication across the observation window eliminates re-scraped identical records.
4
Bot-block detectionIdentifies captcha walls, Cloudflare challenges, and redirect loops so bad scrapes are flagged, not published.
5
Encoding / format validationMalformed JSON, truncated HTML, and non-UTF-8 responses are quarantined before normalisation.
B
Layer BContent Extraction
8 checks
1
Entity recognition accuracyNamed-entity tagging confirms the insight is actually about the tracked competitor, not a partner or acquirer.
2
Date extraction confidenceStructured date fields are verified against publication metadata to catch mis-scraped timestamps.
3
Quote / claim isolationVerbatim strings are kept separate from paraphrase; only verbatim strings can support a Confirmed claim.
4
Competitor name disambiguationApollo (CRM) vs Apollo (other) style conflicts are resolved by URL + sector context before storage.
5
Product vs company distinctionSignals about a product line are not promoted to company-level findings without corroboration.
6
Pricing normalisationCurrencies, billing periods, and per-seat vs flat-rate are standardised before any comparison is made.
7
Geographic scope taggingCountry-specific pricing or availability is tagged so it is not presented as a global claim.
8
Language detectionNon-English content is flagged and excluded unless a validated translation is available.
C
Layer CInsight Generation
10 checks
1
Hallucination detectionCross-source agreement check: a claim is only Confirmed if two independent source types corroborate it.
2
Citation verificationEvery claim must trace back to a stored source URL with a scrape timestamp. No source means the claim is dropped.
3
Scope inflation detectionCatches insights that overreach their evidence — e.g., "dominating the market" from one review score.
4
Sentiment calibrationRaw sentiment scores are adjusted against category baseline to avoid over-indexing on outlier reviews.
5
Competitive framing accuracyRelative claims ("cheaper than", "faster than") require both sides of the comparison to be in evidence.
6
Actionability scoreInsights are scored 1–5 on whether a marketing, sales, or product team can act on them. Sub-3 insights are dropped.
7
Novelty checkInsights are compared against the last 90 days of published findings; re-surfaced stale signals are flagged.
8
Keyword stuffing detectionFlags insights that repeat keyword phrases without meaningful variance — a signal of synthesis failure.
9
Internal consistencyIf the same competitor appears in two insights with contradictory claims, both are held for human review.
10
Character limit complianceInsight length bounds are enforced so briefings stay scannable; over-long synthesis is re-run, not truncated.
D
Layer DQA & Dispatch
7 checks
1
Confidence threshold gateOnly insights scoring ≥85% on the composite confidence model make it into customer briefings.
2
Human-review flag triggerAnything near the 85% boundary, or touching sensitive claims, is auto-flagged for a second-pass review.
3
Duplicate insight detectionA final dedup pass across all verified insights before digest assembly prevents customers receiving the same signal twice.
4
Customer relevance scoringEach insight is scored against the customer's configured competitor set and market context before inclusion.
5
Digest assembly validationStructure, section order, and required fields are validated against the digest schema before rendering.
6
ISP / spam header checksEmail delivery headers are verified against SPF/DKIM/DMARC records to maximise inbox placement.
7
Delivery receipt loggingEvery sent briefing writes a delivery record; bounces and failed sends trigger an immediate alert.

How the 94% Confidence Score Is Calculated

Confidence is a composite, not a single metric. Every insight receives four independent sub-scores, then a weighted average is computed. The 94% figure is the median across all insights shipped in the last 90 days that received customer feedback.

Source authority30%

Primary-source URLs (official pricing page, official press release, ad library record) score highest. Aggregator mentions, second-hand blog coverage, and community posts score lower. The score is pre-assigned per source type and cannot be negotiated up by volume.

Multi-source agreement35%

The single most important signal. A claim confirmed by two or more independent source types (e.g. a job posting plus a changelog entry) scores near-maximum. A claim from a single source caps at 0.72 on this dimension regardless of other factors.

Citation completeness20%

Does the insight link to a stored, timestamped source URL? Does it include the verbatim string it was derived from? Derived figures must also carry a recorded calculation. Missing any element drops this sub-score to zero, which floors the overall confidence below the 85% dispatch gate.

Claim scope fit15%

Measured by the scope inflation check (Layer C, check 3). Does the claim stay within what the evidence actually shows? A specific tier launch from a pricing page update is well-scoped. A broad market claim from a single LinkedIn post is over-scoped and loses points here.

The dispatch gate

Only insights scoring ≥85% on the composite confidence model enter customer briefings. Insights between 75–85% are flagged for human review. Anything below 75% is dropped and logged for pipeline diagnostics. The 85% gate is why the median of shipped insights lands at 94% — most pass with margin.

The 98% Accuracy Commitment

Accuracy is the percentage of shipped claims that have not been contested or corrected by customers, competitors, or our own internal QA cycle. As of Q2 2026, fewer than 2% of shipped claims have resulted in a correction request. We disclose every correction publicly on the relevant report page.

The biggest failure mode in AI-generated competitive intelligence is hallucinated figures. In early 2026 we shipped a battlecard claiming a competitor was undercutting a rival by over 60% — the actual figure was 33%, baselined against the wrong price. That correction prompted every derived figure to now carry a mandatory audit trail:

Required fields for any derived figure
The claim itself
The baseline value used
The source the baseline came from (URL + scrape date)
The exact calculation
The verifier's confidence rating for that derivation

You see this audit trail in the “View evidence” expand panel under any Key Finding that contains a derived figure. If a derivation looks wrong, the audit trail tells you exactly what to check.

Monitoring Cadence

Scrape run
Every Sunday UTC

The full 350+ source pipeline runs once a week. Each competitor is re-scraped end-to-end.

Briefing delivery
Every Monday

Verified insights from the Sunday scrape land in customer inboxes by Monday morning.

Max observation age
9 days

A report read on Tuesday cites observations at most nine days old. No stale signals.

Correction cycle
5 business days

Any contested claim receives one of three responses (update, retract, or stand by) within five business days.

IndustryLens is weekly, not real-time. A competitive move that happens on a Tuesday will land in reports the following Monday, not within hours. We are transparent about this limitation. If your team needs hour-grain signals, IndustryLens is not the right tool.

Frequently Asked Questions

Why 350+ sources — isn't that duplicative?

Source overlap is a feature, not a bug. When G2, Glassdoor, and a company blog all independently describe the same feature launch, the confidence score rises. When they contradict each other, a flag is raised instead of an insight being shipped. Redundancy is how hallucination detection works at scale.

What does "94% confidence" actually mean?

The 94% figure is the median confidence score across all insights shipped in the past 90 days that have received user feedback. Confidence is calculated as a weighted composite of: source authority, multi-source agreement, claim scope fit, and citation completeness. Only insights above the 85% gate enter customer briefings — so the 94% median reflects that most pass with margin to spare.

How is "98% accuracy" defined and measured?

Accuracy is the percentage of shipped claims that have not been contested or corrected by customers, competitors, or our own internal QA cycle. As of Q2 2026, fewer than 2% of shipped claims have resulted in a correction request. We disclose every correction publicly on the relevant report page.

Do you use AI to generate insights?

Yes — an AI synthesis pass converts raw observations into insight drafts, but that pass is followed by an adversarial verification pass that checks every claim against its source. Insights that fail verification are dropped, not edited. The pipeline also runs 30 automated checks across four layers before anything reaches a customer briefing.

How does IndustryLens differ from Klue or Crayon?

Klue and Crayon are strong enterprise tools that require a dedicated CI analyst to get value. IndustryLens is built for teams without that resource — one person can run CI for the entire org. We also differ on transparency: every claim ships with a confidence score and a clickable source URL, so your team can verify what we deliver.

What happens if a source goes behind a login or changes its format?

Bot-block detection (Layer A, check 4) and format validation (Layer A, check 5) catch this automatically. The source is flagged as degraded, removed from the active pool for that run, and a pipeline alert is triggered. Insights that relied on that source in previous runs are not re-surfaced until the source recovers.

Corrections and Disputes

If you see a derivation, quote, or finding you believe is wrong: contact us with the report URL and the specific finding. You will get one of three responses within five business days:

1
Update with citationWe agree — the report is corrected and you are credited in the dateModified note.
2
RetractWe agree the finding should not have shipped — we pull it and explain why on the report page.
3
Stand by it with explanationWe disagree — we lay out the full audit trail so you can judge.

We do not silently edit reports. Every change appears in the dateModified field.

See the pipeline working on your competitors

Start a 30-day trial. Add your competitors and get your first verified Monday briefing — with confidence scores, source URLs, and audit trails — within a week.

No credit card · 30 days free · See pricing · About IndustryLens