Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains

Fonseca, Marcio; Cohen, Shay B.

Computer Science > Computation and Language

arXiv:2311.08704 (cs)

[Submitted on 15 Nov 2023 (v1), last revised 27 Jun 2024 (this version, v2)]

Title:Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains

Authors:Marcio Fonseca, Shay B. Cohen

View PDF HTML (experimental)

Abstract:Although large language models (LLMs) exhibit remarkable capacity to leverage in-context demonstrations, it is still unclear to what extent they can learn new concepts or facts from ground-truth labels. To address this question, we examine the capacity of instruction-tuned LLMs to follow in-context concept guidelines for sentence labeling tasks. We design guidelines that present different types of factual and counterfactual concept definitions, which are used as prompts for zero-shot sentence classification tasks. Our results show that although concept definitions consistently help in task performance, only the larger models (with 70B parameters or more) have limited ability to work under counterfactual contexts. Importantly, only proprietary models such as GPT-3.5 and GPT-4 can recognize nonsensical guidelines, which we hypothesize is due to more sophisticated alignment methods. Finally, we find that Falcon-180B-chat is outperformed by Llama-2-70B-chat is most cases, which indicates that careful fine-tuning is more effective than increasing model scale. Altogether, our simple evaluation method reveals significant gaps in concept understanding between the most capable open-source language models and the leading proprietary APIs.

Comments:	ACL 2024 camera ready
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2311.08704 [cs.CL]
	(or arXiv:2311.08704v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2311.08704

Submission history

From: Marcio Fonseca [view email]
[v1] Wed, 15 Nov 2023 05:11:26 UTC (240 KB)
[v2] Thu, 27 Jun 2024 03:48:35 UTC (246 KB)

Computer Science > Computation and Language

Title:Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators