English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Deep Dive: Revisiting Prompt Engineering — A Comprehensive Evaluation for LLM-based Personalized Recommendation

Forum topic · ✨步子哥 · 2025-12-11

Summary

A forum post analyzes the paper 'Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation' by Kusano et al., which conducts the largest empirical study to date on prompt engineering for LLM-driven single-user personalized recommendation. The study evaluates 23 prompt types across 8 public datasets and 12 LLMs (cost-efficient and high-performance), using statistical tests and linear mixed-effects models to assess recommendation accuracy (e.g., nDCG) and inference cost. Key findings: for cost-efficient LLMs, three prompt strategies help most—rephrasing instructions, incorporating background knowledge, and simplifying reasoning steps; for high-performance LLMs, simple prompts outperform complex ones, achieving higher accuracy at lower cost; popular NLP techniques like step-by-step reasoning prompts and dedicated reasoning models degrade recommendation accuracy. The article distills practical guidance: match prompt complexity to model capability, balance accuracy against inference cost, avoid blindly reusing NLP prompt tricks, and iterate prompt choices empirically. It concludes that prompt design is critical in privacy-sensitive, single-user recommendation settings and offers evidence-based best practices for deploying LLM recommenders.

This post is a deep analysis of the paper *Revisiting Prompt Engineering: A Comprehensive Evaluation for LLM-based Personalized Recommendation* (Kusano et al.), originally published on zhichai.net.

Introduction

Large language models (LLMs) make it possible to perform recommendation tasks via natural language prompts. Compared with traditional collaborative filtering, LLM-based recommendation shows unique advantages in cold-start, cross-domain recommendation, and zero-shot scenarios, supports flexible input formats, and can generate explanations for user behavior. However, how to design prompts to fully unlock LLMs' recommendation potential has lacked systematic study. This paper fills that gap with a large-scale evaluation.

The paper focuses on single-user personalized recommendation—using only the target user's own history, without other users' data. This setting matters for privacy-sensitive or data-limited applications, where prompt engineering becomes the key lever for controlling output quality. The authors compare 23 prompt types across 8 public datasets and 12 LLMs, evaluating accuracy and inference cost via statistical tests and linear mixed-effects models—the most comprehensive empirical analysis of prompt engineering for LLM recommendation to date.

Advantages and Challenges of LLM-based Recommendation

Advantages:

  • Handling cold start: LLMs can reason from prior knowledge when interaction data is missing.
  • Cross-domain recommendation via general semantic understanding.
  • Flexible natural language input, integrating reviews, descriptions, and other unstructured context.
  • Generating explanations, improving user trust.
  • Challenges: In the single-user setting, the model must infer preferences from limited history alone, making prompt design critical. Different LLMs—cost-efficient vs. high-performance—respond very differently to the same prompts, so balancing accuracy and inference cost is essential.

    Experimental Design

  • 23 prompt templates spanning standardized phrases, non-conversational, and conversational prompts, including instruction rephrasing, background-knowledge injection, and step-by-step (chain-of-thought-style) reasoning.
  • 8 public datasets across domains (movies, products, music) to verify generalizability.
  • 12 LLMs, both cost-efficient and high-performance.
  • Metrics: ranking accuracy (e.g., nDCG) and inference cost (e.g., cost of processing 1,600 users per prompt–model combination).
  • Analysis: statistical significance tests and linear mixed-effects models to quantify prompt effects while accounting for dataset and model differences.
  • Key Findings

  • For cost-efficient LLMs, three prompt families are especially effective:
  • 1. Rephrasing instructions for clarity (e.g., rewriting 'recommend movies' as a more specific instruction). 2. Incorporating background knowledge (item attributes, user-preference summaries) as extra cues. 3. Simplifying reasoning—clearer, shorter logic chains that small models can follow reliably.
  • For high-performance LLMs, simple prompts beat complex ones. Verbose or elaborate prompts can reduce accuracy and add unnecessary cost; concise prompts let large models exploit their internal knowledge while cutting inference overhead. In short: *less is more* when the model is capable.
  • Popular NLP prompt tricks don't transfer: step-by-step reasoning prompts lower recommendation accuracy, and dedicated reasoning models also underperform. Recommendation relies more on matching user preference to item attributes than on explicit logical deduction.
  • Accuracy–cost trade-off: higher accuracy usually costs more, but not always—for high-performance models, simple prompts maintain high accuracy while reducing cost, offering better price-performance for deployment.
  • Practical Guidance

  • Match prompt complexity to model capability: use carefully engineered complex prompts (rephrase instructions, add knowledge, simplify reasoning) for small models; use simple, clear prompts for large models.
  • Balance accuracy and cost: prioritize high-performance models + tuned prompts when accuracy matters most; favor cost-efficient models with validated prompts, or simple prompts on large models, when latency/budget matters.
  • Don't blindly reuse NLP prompt techniques: design prompts around user–item matching rather than step-by-step logic.
  • Keep iterating: test multiple prompts on your user base and adapt as models and techniques evolve.

Conclusion

The paper demonstrates that prompt engineering is crucial in single-user LLM recommendation, with effects that are significant and complex. Well-designed prompts greatly boost small models, while simple prompts are both effective and economical for large models. Common NLP prompting methods often fail in recommendation. These findings give developers concrete, evidence-based guidelines for choosing prompts and models—and point toward future work in automated prompt selection and multimodal prompt design for smarter, more efficient, and more trustworthy recommenders.

Tags

#llm#prompt-engineering#recommender-systems#personalized-recommendation#paper-analysis#evaluation#ndcg#inference-cost

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/176415116