Evaluation of Siamese Networks for Semantic Code Search

Sinha, Raunak; Desai, Utkarsh; Tamilselvam, Srikanth; Mani, Senthil

Computer Science > Software Engineering

arXiv:2011.01043 (cs)

[Submitted on 12 Oct 2020]

Title:Evaluation of Siamese Networks for Semantic Code Search

Authors:Raunak Sinha, Utkarsh Desai, Srikanth Tamilselvam, Senthil Mani

View PDF

Abstract:With the increase in the number of open repositories and discussion forums, the use of natural language for semantic code search has become increasingly common. The accuracy of the results returned by such systems, however, can be low due to 1) limited shared vocabulary between code and user query and 2) inadequate semantic understanding of user query and its relation to code syntax. Siamese networks are well suited to learning such joint relations between data, but have not been explored in the context of code search. In this work, we evaluate Siamese networks for this task by exploring multiple extraction network architectures. These networks independently process code and text descriptions before passing them to a Siamese network to learn embeddings in a common space. We experiment on two different datasets and discover that Siamese networks can act as strong regularizers on networks that extract rich information from code and text, which in turn helps achieve impressive performance on code search beating previous baselines on $2$ programming languages. We also analyze the embedding space of these networks and provide directions to fully leverage the power of Siamese networks for semantic code search.

Subjects:	Software Engineering (cs.SE); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2011.01043 [cs.SE]
	(or arXiv:2011.01043v1 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2011.01043

Submission history

From: Utkarsh Desai [view email]
[v1] Mon, 12 Oct 2020 06:07:39 UTC (18,862 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.SE

< prev | next >

new | recent | 2020-11

Change to browse by:

cs
cs.AI
cs.CL

References & Citations

DBLP - CS Bibliography

listing | bibtex

Srikanth Tamilselvam
Senthil Mani

export BibTeX citation

Computer Science > Software Engineering

Title:Evaluation of Siamese Networks for Semantic Code Search

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:Evaluation of Siamese Networks for Semantic Code Search

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators