LLM-Quality-Score

Think before you prompt

The LLM-Quality-Score is a digital nudge that evaluates the quality of a prompt before it is submitted. Based on selected criteria such as clarity, structure, context depth, language quality, bias risk and format requirements, users receive direct feedback on their input and are thereby encouraged to reflect on their prompt once more and to improve it where necessary. The nudge works not through coercion or blocking, but through visible, immediate feedback that directs attention to one’s own input, fully in the spirit of the principle „Think before you prompt.“ In terms of implementation, the score first recognizes the type of prompt via a meta level, for example explaining, summarizing, programming or analyzing, and weights the criteria differently depending on the type. On this basis it calculates a quality value and displays it directly in the input field. In addition, the system stores concrete improvement suggestions following the pattern „Instead of … better …“, so that users receive not only an assessment but also actionable guidance. The underlying principle follows the chain: prompt, quality score, reflection, improved prompt formulation and finally a more accurate AI response.

Accept the cookie setting to see Panopto content.

Unlock Panopto

What does the topic mean?

Large Language Models (LLMs) such as ChatGPT are increasingly being used in higher education for research and learning processes. Our concept is based on Shaw and Nave’s Tri-System Theory, which extends the classical model of human cognitive processes by incorporating artificial intelligence. It describes three interconnected systems: System 1 represents fast and intuitive thinking, System 2 represents deliberate, critical, and analytical thinking, and System 3 represents artificial intelligence, which provides information and suggested solutions. Ideally, System 3 supports the human cognitive systems, while System 2 critically evaluates the AI’s outputs. In practice, however, this verification step is often skipped, as users tend to trust AI-generated responses immediately. This behavior is referred to as cognitive surrender. It should be distinguished from cognitive offloading, where AI is deliberately used as a support tool while the user’s own thinking remains active. Since students often formulate their prompts spontaneously and imprecisely, the resulting responses are frequently incomplete or incorrect and are then adopted uncritically. This is precisely where our nudge comes in.

Goal of the nudge

The goal of the LLM-Quality-Score is to encourage users to consciously reflect on their input before they submit it. By triggering a brief, examining pause, the score is intended to reactivate the reflective thinking of System 2 and thus reduce the tendency toward cognitive surrender. In this way, an AI’s responses are adopted less blindly and language models are used more consciously overall. Shaw and Nave show empirically that feedback and incentives can reactivate the thinking of System 2 and measurably reduce cognitive surrender, which supports the logic behind our score. In the long term, this should help users develop a lasting competence in designing prompts. On a broader level, we pursue the goal of integrating the LLM-Quality-Score as an upstream reflection mechanism into existing artificial intelligence systems at universities.

    Needs analysis

    • Prevalence of language models in studies
      LLMs have become a standard tool for research and learning in everyday university life, so that their use is hardly questioned consciously anymore.
    • Uncritical adoption of outputs
      Shaw and Nave show that people follow an AI’s recommendations in roughly four out of five cases even when these are faulty, and that their accuracy rises or falls directly with the quality of the output.
    • Deceptive confidence
      Access to an AI increases confidence in one’s own answer even when it is wrong, so that errors go unnoticed.
    • Low prompt competence
      Spontaneous and imprecise prompts frequently lead to incomplete or faulty answers without the cause being recognized.
    • Missing impulse to reflect
      In existing systems there is so far no mechanism that encourages thinking about one’s own input before submitting it.

    Cause analysis

    • Automation bias

      The tendency to trust an AI’s outputs more than one’s own judgment leads to answers being adopted without scrutiny.

    • Cognitive surrender (after Shaw and Nave)

      The examining thought of System 2 is skipped and the output of System 3 is uncritically accepted as one’s own answer.

    • Processing fluency

      Answers phrased fluently and confidently appear more convincing and are more readily taken to be correct, regardless of their actual accuracy.

    • Authority heuristic

      The AI is perceived as a competent and neutral authority, so that its outputs are granted high credibility and questioned less often.

    • Principle of least effort (cognitive miser)

      Because reflective thinking is effortful, users choose the convenient path and adopt the quickly available answer.

Target Group

The primary target group consists of university computing centers and external institutions from research and technology such as RWTH Aachen, HAWK or the KI-Campus. They develop and operate the universities’ own language models, for example HSHL with KI:Connect, and are therefore the right place to integrate the LLM-Quality-Score as an upstream reflection mechanism. Indirectly, the nudge addresses the students who use these systems every day.

Added value of the nudge

  • Better answers through better prompts

    A brief impulse to reflect leads to more precise input and thus to more accurate and useful AI responses.

  • Protection against uncritical adoption

    The score reactivates System 2 and reduces the risk of adopting faulty outputs without scrutiny.

  • Lasting prompt competence

    Over time, users learn to deal with AI more consciously, more critically and more competently.

CONTACT US

Long Phi Do

long-phi.do@stud.hshl.de

David Klein

david.klein@stud.hshl.de

Nethen-Kiya Rosentreter

nethen-kiya.rosentreter@stud.hshl.de

Vasilios Kiosses

vasilios.kiosses@stud.hshl.de