Purpose of this tool: Compare two prompt variants and their corresponding LLM outputs directly. Enter the texts, set the weight and score (1-5) for each criterion, and read the weighted final result right away.
Variant A
Output Statistics (Estimate):
- Characters: 0
- Words: 0
- Tokens (rough estimate): 0
Variant B
Output Statistics (Estimate):
- Characters: 0
- Words: 0
- Tokens (rough estimate): 0
Evaluation Criteria
Weighted Final Result
Score Variant A
0.00 / 5
Difference (Absolute Delta)
0.00
Score Variant B
0.00 / 5
How to Use This Tool Wisely
Systematically comparing prompt variants is essential for building reliable LLM applications. To avoid making decisions based on chance or personal preferences, a structured approach is necessary. Learn more about the fundamentals in our guide on A/B testing prompts.
Keep the following guidelines in mind while evaluating:
- Isolate variables: Change only one part of the prompt per test (for example, adding an explicit format or a role instruction). Do not change the temperature or model type between two runs.
- Limit anchoring bias with blind scoring: Use the built-in 'Blind mode'. If you know which prompt is yours or which method you prefer, you will unconsciously judge the output with bias.
- Define criteria in advance: Determine what is important before you read the results. Is factual accuracy decisive, or is it really about writing style? Assign the appropriate weights using an established self-evaluation framework.
- Run multiple runs: LLMs are stochastic. A single generation can be lucky or unlucky. For proper quality assurance, it is wise to also combine human evaluation with larger test sets.
Want to discuss your test results or evaluation methodologies? Join the conversation in our LLMnet Community.


