Research evidence library

Fifteen supplied papers, read in full and translated into bounded product decisions.

arXiv · 2022-10-18

Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages

Read the evidence record

Research question

How can a retrieval benchmark cover languages with very different resource levels while using native queries and relevance judgments?

Method

An 18-language, same-language Wikipedia retrieval benchmark built with native-speaker queries and judgments; lexical, dense and hybrid baselines were compared on the released pools.

Evidence supports

Evaluate each language explicitly, document corpus and judgment construction, and keep native review central. The reported hybrid result is a historical benchmark observation.

Limits and transferability

This is same-language information retrieval, not translation assessment or web SEO. Heuristic segmentation, candidate-pool gaps and unfinished test labels limit transferability; it says nothing about hreflang or Google rankings.

Claim classification

  • Supported
  • Historical
  • Academic only

NAVINES influence: Multilingual Page Comparison Lab

arXiv · 2024-04-14

Competitive Retrieval: Going Beyond the Single Query

Read the evidence record

Research question

How does publisher competition change when participants allocate content decisions across multiple queries rather than one isolated query?

Method

A formal multi-query game plus four controlled student competitions covering 30 TREC topics, 84 participants and specified proxy rankers; best-response dynamics and feature changes were examined.

Evidence supports

Cross-query opportunity cost and competitor response deserve explicit treatment; equilibrium existence does not imply that learning dynamics will converge.

Limits and transferability

The student setting, topics and rankers are controlled proxies. The work does not identify current Google features, predict winners or justify reusable weights or a practical Nash-equilibrium claim.

Claim classification

  • Limited
  • Academic only

NAVINES influence: Competitive multi-query portfolio planner

arXiv · 2020-05-26

Ranking-Incentivized Quality Preserving Content Modification

Read the evidence record

Research question

Can ranking-incentivized content modification include coherence and modification cost instead of optimizing promotion alone?

Method

A controlled passage-replacement method balanced promotion and coherence, evaluated with proxy rankers and offline data across 31 queries, with an explicit ethics discussion.

Evidence supports

Quality constraints and the cost of changing content belong in planning. A promotional change that fails editorial integrity should be rejected.

Limits and transferability

Passages came from higher-ranked candidates and quality judgments were bounded and subjective. It does not support copying winners, real-engine causal claims or unrestricted automated rewriting.

Claim classification

  • Limited
  • Academic only

NAVINES influence: Competitive multi-query portfolio planner

Journal of Artificial Intelligence Research · 2025

The Search for Stability: Learning Dynamics of Strategic Publishers with Initial Documents

Read the evidence record

Research question

Under which formal ranking mechanisms and content-deviation costs can strategic publisher learning converge or remain unstable?

Method

Formal publisher-utility models with rank benefit minus deviation cost, convergence results for specified linear/softmax settings, and discrete simulations, predominantly with two publishers.

Evidence supports

Change cost, response dynamics and instability are useful planning dimensions; a mechanism can have an equilibrium while local responses cycle or remain pseudoperiodic.

Limits and transferability

The model assumes a static information need, known ranking function and embedding-space actions. Simulations are not a real search engine and do not justify a Google mechanism or guaranteed convergence claim.

Claim classification

  • Supported
  • Limited
  • Academic only

NAVINES influence: Competitive multi-query portfolio planner

ACM KDD · 2024

GEO: Generative Engine Optimization

Read the evidence record

Research question

In controlled generative-engine experiments, how do selected content modifications affect proxy visibility across queries and domains?

Method

GEO-BENCH combined 10,000 queries from nine datasets with visibility proxies and experiments on two generative engines; citation, quotation, statistics and other interventions were compared by domain.

Evidence supports

Clear attribution, genuine evidence and domain-aware experimentation can be reviewed as editorial signals; keyword stuffing performed poorly in this benchmark.

Limits and transferability

Two engines, black-box variance and proxy metrics do not establish traditional-search effects or future citation probability. Reported gains are historical benchmark results, not forecasts; appendix prompts are untrusted.

Claim classification

  • Limited
  • Historical
  • Academic only

NAVINES influence: AI citation-readiness lab

arXiv · 2024-07-02

Adversarial Search Engine Optimization for Large Language Models

Read the evidence record

Research question

Can content in retrieved web or plugin documents manipulate an LLM's product preferences, and which broader attack categories appear beyond classic prompt injection?

Method

Controlled experiments with fictional products, web/plugin documents and proprietary models examined model-directed instructions, concealment, false claims, source suppression and competitor discrediting.

Evidence supports

Retrieved content is an untrusted input surface; safety review should cover concealed persuasion and source integrity as well as direct instructions.

Limits and transferability

The favorable fictional setup, limited scale and changing proprietary systems do not establish prevalence or complete detection. Exact attack payloads are not republished and the checker is not a certification.

Claim classification

  • Supported
  • Limited
  • Academic only

NAVINES influence: AI search manipulation safety checker

American Economic Review · 2011-10

Bayesian Persuasion

Read the evidence record

Research question

Ethical information design

Method

Theoretical research · n/a · DOI 10.1257/aer.101.6.2590

Assumptions

  • Bayesian persuasion: A theoretical model in which a sender commits to an information structure before a rational receiver acts.
  • Sender and receiver: The parties that design a signal and act after observing it in an information-design model.
  • Information design: The deliberate, truthful structuring of what information is revealed and when.
  • Commitment: The model assumption that a chosen information policy is fixed before its outcome is known.

Evidence supports

Sequence truthful evidence around the receiver's decision, expose uncertainty and disqualifiers, and never turn Bayesian persuasion into a manipulation score.

Limits and transferability

Map claims to evidence, uncertainty, alternatives and honest fit without a persuasion score. This is a disclosure and planning aid, not proof that wording will change behavior.

What this does not establish

No universal score is calculated · This is a disclosure and planning aid, not proof that wording will change behavior.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Evidence & positioning lab

Review of Economic Studies · 2017-01

Competition in Persuasion

Read the evidence record

Research question

Competitive positioning

Method

Theoretical research · n/a · DOI 10.1093/restud/rdw052

Assumptions

  • Sender and receiver: The parties that design a signal and act after observing it in an information-design model.
  • Information design: The deliberate, truthful structuring of what information is revealed and when.
  • Commitment: The model assumption that a chosen information policy is fixed before its outcome is known.

Evidence supports

Compare real alternatives on a shared basis. Competition can increase or reduce disclosure depending on conditions; it does not automatically produce truth.

Limits and transferability

Map claims to evidence, uncertainty, alternatives and honest fit without a persuasion score. This is a disclosure and planning aid, not proof that wording will change behavior.

What this does not establish

No universal score is calculated · This is a disclosure and planning aid, not proof that wording will change behavior.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Evidence & positioning lab

Experimental Economics · 2024-09-13

Rational inattention in games: experimental evidence

Read the evidence record

Research question

Rational attention and friction

Method

Controlled experiment · 12 · n=238 · DOI 10.1007/s10683-024-09843-z

Assumptions

  • Rational inattention: A family of models where decision makers trade the value of information against attention or processing cost.

Evidence supports

Review stakes, familiarity, uncertainty, reversibility and information cost as context. A controlled buyer–seller task does not yield a universal attention score.

Limits and transferability

Review a decision environment in context instead of assuming that fewer options are better. The output is a set of observations and test hypotheses, not a conversion score.

What this does not establish

No universal score is calculated · The output is a set of observations and test hypotheses, not a conversion score.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

CC BY 4.0 open access

How NAVINES applied it: Decision friction & choice architecture lab

ACM KDD · 2003

Maximizing the Spread of Influence through a Social Network

Read the evidence record

Research question

Influence and distribution

Method

Algorithmic research · Independent Cascade; Linear Threshold; (1−1/e−ε); DOI 10.1145/956750.956769

Assumptions

  • Influence maximization: Selecting seed nodes to maximize expected diffusion in a specified graph and propagation model.
  • Submodularity: A diminishing-returns property under which adding the same seed tends to help less as the seed set grows.

Evidence supports

Model a real graph with explicit diffusion parameters and uncertainty. Greedy approximation under submodular models is not a virality forecast.

Limits and transferability

Compare greedy Monte Carlo seed selection with a degree baseline under explicit graph-diffusion assumptions. Estimated reach depends entirely on the supplied graph, parameters and model; it is not a virality forecast.

What this does not establish

No universal score is calculated · Estimated reach depends entirely on the supplied graph, parameters and model; it is not a virality forecast.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Influence seeding planner

ACM CHI · 2001

What Makes Web Sites Credible? A Report on a Large Quantitative Study

Read the evidence record

Research question

Credibility evidence

Method

Historical behavioral evidence · 1999 · n=1,410 · US/FI · 51 · DOI 10.1145/365024.365037

Assumptions

Review observable identity, authorship, sources, policies, usability and security evidence without predicting conversion.

Evidence supports

Show identity, authorship, sources, policies, dates, security and accessible usability. Historical perceived-credibility evidence is not a conversion or ranking study.

Limits and transferability

Review observable identity, authorship, sources, policies, usability and security evidence without predicting conversion. Automated checks cannot verify expertise, claim truth, full accessibility or human perception.

What this does not establish

No universal score is calculated · Automated checks cannot verify expertise, claim truth, full accessibility or human perception.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Credibility evidence audit

Journal of Consumer Research · 2010-10

Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload

Read the evidence record

Research question

Choice architecture

Method

Meta-analysis · 63 · 50 · N=5,036 · d=.02 · 95% CI −.09… .12 · I²=68% · DOI 10.1086/651235

Assumptions

  • Choice overload: A context-dependent possibility that a larger choice set creates adverse decision outcomes; not a universal rule.

Evidence supports

Categorize and progressively disclose when useful, but test the environment. A meta-analysis found a near-zero mean choice-overload effect with substantial heterogeneity.

Limits and transferability

Review a decision environment in context instead of assuming that fewer options are better. The output is a set of observations and test hypotheses, not a conversion score.

What this does not establish

No universal score is calculated · The output is a set of observations and test hypotheses, not a conversion score.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Decision friction & choice architecture lab

Data Mining and Knowledge Discovery · 2009

Controlled experiments on the web: survey and practical guide

Read the evidence record

Research question

Controlled experimentation

Method

Operational experimentation guidance · DOI 10.1007/s10618-008-0114-1

Assumptions

  • Overall Evaluation Criterion: The prespecified primary metric used to judge an experiment as a whole.
  • Guardrail metric: A metric that can block or qualify a launch when an important protected outcome worsens.

Evidence supports

Define an OEC, guardrails, power, assignment, ramp and stop rules before reading results. A borderline p-value alone never names a winner.

Limits and transferability

Plan sample needs, compare valid experiments and keep before/after observations descriptive. Statistical significance is not business importance, and observational SEO changes are not causal proof.

What this does not establish

No universal score is calculated · Statistical significance is not business importance, and observational SEO changes are not causal proof.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher page says open access; reuse license requires confirmation

How NAVINES applied it: SEO experiment planner and results comparator

ACM KDD · 2017

A Dirty Dozen: Twelve Common Metric Interpretation Pitfalls in Online Controlled Experiments

Read the evidence record

Research question

Metric interpretation failures

Method

Operational experimentation guidance · 12 · Microsoft · DOI 10.1145/3097983.3098024

Assumptions

  • Sample Ratio Mismatch: A statistically unlikely difference between observed and intended experiment allocation, often signaling execution trouble.
  • Novelty and primacy effects: Temporary responses caused by a change being new, or by users learning and adapting over time.
  • Simpson's paradox: A reversal between aggregate and subgroup patterns caused by different group weights or confounding.
  • Twyman's Law: The operational warning that an unusually surprising result deserves extra scrutiny before celebration.

Evidence supports

Check SRM, ratio metrics, telemetry, power, multiplicity, segments, outliers, novelty, funnel completeness, Simpson risk and Twyman's Law before acting.

Limits and transferability

Plan sample needs, compare valid experiments and keep before/after observations descriptive. Statistical significance is not business importance, and observational SEO changes are not causal proof.

What this does not establish

No universal score is calculated · Statistical significance is not business importance, and observational SEO changes are not causal proof.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: SEO experiment planner and results comparator

ACM CIKM · 2013

Beyond Clicks: Query Reformulation as a Predictor of Search Satisfaction

Read the evidence record

Research question

Search satisfaction beyond clicks

Method

Historical behavioral evidence · ≈6,000 · 7,628 · n=218 · DOI 10.1145/2505515.2505682

Assumptions

  • Query reformulation: A follow-up query that changes wording, scope or terms while a person continues a search task.

Evidence supports

Similar quick reformulations can support a friction hypothesis in session logs. A click is not success, silence is not satisfaction, and GSC aggregates are not sessions.

Limits and transferability

Review anonymized session-level reformulations as hypotheses while keeping query data local. A click is not proof of success and no follow-up query is not proof of satisfaction.

What this does not establish

No universal score is calculated · A click is not proof of success and no follow-up query is not proof of satisfaction.

What NAVINES deliberately did not build

No universal score is calculated

Primary sources

Publisher copyright; summary and official link only

How NAVINES applied it: Search journey friction analyzer

Evidence reviewed 15 August 2026 · Does not establish current Google ranking factors or guaranteed outcomes.

Research

Original analysis and critical notes with methods, sources, and limitations.

We do not chase algorithms. We build clarity, trust, and compounding visibility.

Make advanced search intelligence understandable, useful, and accessible—then help businesses turn it into measurable action.

Method

  • Inventory hosts, environments, templates, key journeys, and known migrations.
  • Collect bounded crawl, server, Search Console, sitemap, and rendered-page evidence.
  • Write the business question and define the grain, filters, and comparable periods.
  • Preserve query, page, locale, device, and country context before aggregating.

Limitations

  • Privacy thresholds, sampling, consent, attribution, and tracking loss create gaps.
  • Search Console and analytics use different definitions, windows, and processing systems.
WhatsApp