Zum Hauptinhalt springen

Showing 1–6 of 6 results for author: Drosos, I

Searching in archive cs. Search in all archives.
.
  1. arXiv:2408.08781  [pdf, other

    cs.AI cs.CL

    Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions

    Authors: Bhuvanashree Murugadoss, Christian Poelitz, Ian Drosos, Vu Le, Nick McKenna, Carina Suzana Negreanu, Chris Parnin, Advait Sarkar

    Abstract: LLMs-as-a-judge is a recently popularized method which replaces human judgements in task evaluation (Zheng et al. 2024) with automatic evaluation using LLMs. Due to widespread use of RLHF (Reinforcement Learning from Human Feedback), state-of-the-art LLMs like GPT4 and Llama3 are expected to have strong alignment with human preferences when prompted for a quality judgement, such as the coherence o… ▽ More

    Submitted 16 August, 2024; originally announced August 2024.

  2. "It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study

    Authors: Ian Drosos, Advait Sarkar, Xiaotong Xu, Carina Negreanu, Sean Rintel, Lev Tankelevitch

    Abstract: Generative AI tools can help users with many tasks. One such task is data analysis, which is notoriously challenging for non-expert end-users due to its expertise requirements, and where AI holds much potential, such as finding relevant data sources, proposing analysis strategies, and writing analysis code. To understand how data analysis workflows can be assisted or impaired by generative AI, we… ▽ More

    Submitted 3 July, 2024; originally announced July 2024.

    Comments: Ian Drosos, Advait Sarkar, Xiaotong Xu, Carina Negreanu, Sean Rintel, and Lev Tankelevitch. 2024. "It's like a rubber duck that talks back": Understanding Generative AI-Assisted Data Analysis Workflows through a Participatory Prompting Study. In Proceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work (CHIWORK 2024)

    Journal ref: Proceedings of the 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work (CHIWORK 2024)

  3. Improving Steering and Verification in AI-Assisted Data Analysis with Interactive Task Decomposition

    Authors: Majeed Kazemitabaar, Jack Williams, Ian Drosos, Tovi Grossman, Austin Henley, Carina Negreanu, Advait Sarkar

    Abstract: LLM-powered tools like ChatGPT Data Analysis, have the potential to help users tackle the challenging task of data analysis programming, which requires expertise in data processing, programming, and statistics. However, our formative study (n=15) uncovered serious challenges in verifying AI-generated results and steering the AI (i.e., guiding the AI system to produce the desired output). We develo… ▽ More

    Submitted 1 August, 2024; v1 submitted 2 July, 2024; originally announced July 2024.

    Comments: Published at UIST 2024; 19 pages, 9 figures, and 2 tables

    Journal ref: Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (UIST 2024)

  4. arXiv:2404.07114  [pdf, other

    cs.HC

    "My toxic trait is thinking I'll remember this": gaps in the learner experience of video tutorials for feature-rich software

    Authors: Ian Drosos, Advait Sarkar, Andrew D. Gordon

    Abstract: Video tutorials are a popular medium for informal and formal learning. However, when learners attempt to view and follow along with these tutorials, they encounter what we call gaps, that is, issues that can prevent learning. We examine the gaps encountered by users of video tutorials for feature-rich software, such as spreadsheets. We develop a theory and taxonomy of such gaps, identifying how th… ▽ More

    Submitted 10 April, 2024; originally announced April 2024.

  5. arXiv:2312.16633  [pdf, ps, other

    cs.HC

    Participatory prompting: a user-centric research method for eliciting AI assistance opportunities in knowledge workflows

    Authors: Advait Sarkar, Ian Drosos, Rob Deline, Andrew D. Gordon, Carina Negreanu, Sean Rintel, Jack Williams, Benjamin Zorn

    Abstract: Generative AI, such as image generation models and large language models, stands to provide tremendous value to end-user programmers in creative and knowledge workflows. Current research methods struggle to engage end-users in a realistic conversation that balances the actually existing capabilities of generative AI with the open-ended nature of user workflows and the many opportunities for the ap… ▽ More

    Submitted 27 December, 2023; originally announced December 2023.

    Comments: Proceedings of the 34th Annual Conference of the Psychology of Programming Interest Group (PPIG 2023)

    Journal ref: Proceedings of the 34th Annual Conference of the Psychology of Programming Interest Group (PPIG 2023)

  6. arXiv:2310.01297  [pdf, other

    cs.HC cs.AI cs.CL cs.PL

    Co-audit: tools to help humans double-check AI-generated content

    Authors: Andrew D. Gordon, Carina Negreanu, José Cambronero, Rasika Chakravarthy, Ian Drosos, Hao Fang, Bhaskar Mitra, Hannah Richardson, Advait Sarkar, Stephanie Simmons, Jack Williams, Ben Zorn

    Abstract: Users are increasingly being warned to check AI-generated content for correctness. Still, as LLMs (and other generative models) generate more complex output, such as summaries, tables, or code, it becomes harder for the user to audit or evaluate the output for quality or correctness. Hence, we are seeing the emergence of tool-assisted experiences to help the user double-check a piece of AI-generat… ▽ More

    Submitted 2 October, 2023; originally announced October 2023.