Exposing Query Identification for Search Transparency

Li, Ruohan; Li, Jianxiang; Mitra, Bhaskar; Diaz, Fernando; Biega, Asia J.

Computer Science > Information Retrieval

arXiv:2110.07701 (cs)

[Submitted on 14 Oct 2021 (v1), last revised 11 Apr 2022 (this version, v3)]

Title:Exposing Query Identification for Search Transparency

Authors:Ruohan Li, Jianxiang Li, Bhaskar Mitra, Fernando Diaz, Asia J. Biega

View PDF

Abstract:Search systems control the exposure of ranked content to searchers. In many cases, creators value not only the exposure of their content but, moreover, an understanding of the specific searches where the content is surfaced. The problem of identifying which queries expose a given piece of content in the ranking results is an important and relatively under-explored search transparency challenge. Exposing queries are useful for quantifying various issues of search bias, privacy, data protection, security, and search engine optimization.
Exact identification of exposing queries in a given system is computationally expensive, especially in dynamic contexts such as web search. We explore the feasibility of approximate exposing query identification (EQI) as a retrieval task by reversing the role of queries and documents in two classes of search systems: dense dual-encoder models and traditional BM25 models. We then propose how this approach can be improved through metric learning over the retrieval embedding space. We further derive an evaluation metric to measure the quality of a ranking of exposing queries, as well as conducting an empirical analysis focusing on various practical aspects of approximate EQI. Overall, our work contributes a novel conception of transparency in search systems and computational means of achieving it.

Subjects:	Information Retrieval (cs.IR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2110.07701 [cs.IR]
	(or arXiv:2110.07701v3 [cs.IR] for this version)
	https://doi.org/10.48550/arXiv.2110.07701

Submission history

From: Bhaskar Mitra [view email]
[v1] Thu, 14 Oct 2021 20:19:27 UTC (5,678 KB)
[v2] Mon, 17 Jan 2022 14:43:22 UTC (5,682 KB)
[v3] Mon, 11 Apr 2022 14:49:53 UTC (463 KB)

Computer Science > Information Retrieval

Title:Exposing Query Identification for Search Transparency

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Retrieval

Title:Exposing Query Identification for Search Transparency

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators