What Is Prompt Engineering? A Practical Guide for Developers
Prompt engineering is the work of writing and evaluating instructions for language models. Learn a repeatable workflow for clearer prompts, useful examples, and reliable results.
Prompt engineering is the design and testing of the instructions and context you give a language model so its responses meet defined requirements. It is a practical development process, not a search for a magic phrase: outputs can vary, and the same prompt may behave differently across model types or versions.
A reliable starting point is to define what a good answer must do, write a specific prompt with the needed context and output constraints, then evaluate it against representative examples. If results still fail, diagnose the failure before adding more wording: a model capability, latency, cost, or application design issue may call for a different solution.
1. What is prompt engineering?
Prompt engineering is the process of writing effective instructions for a model so it generates content that meets your requirements. In practice, a prompt can include the task, relevant context, constraints, examples, and the desired response format. Prompt design is iterative: you form a hypothesis about what the model needs, test it, inspect failures, and refine.
It is not a guarantee of identical output. Language model generation is non-deterministic, and different model types or snapshots in the same family can respond differently. OpenAI, Google, and Anthropic all frame their prompt guidance as something to test against the actual use case rather than apply as a universal recipe. See the [OpenAI prompt engineering guide](https://developers.openai.com/api/docs/guides/prompt-engineering), [Google prompt design strategies](https://ai.google.dev/gemini-api/docs/prompting-strategies), and [Anthropic prompt engineering overview](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/overview).
2. Start with success criteria
Before editing a prompt, state what success means. A useful specification names the task, required content, unacceptable outcomes, and constraints the response must obey. Also decide how you will test those criteria. This helps distinguish an actual improvement from a response that merely sounds better.
| Criterion | Example for a support-answer generator |
|---|---|
| Task | Answer a customer question using the supplied product documentation. |
| Required content | Give the relevant steps and identify any prerequisite. |
| Must avoid | Do not invent product behavior or claim an action succeeded. |
| Format | Return a short answer followed by numbered steps. |
| Evaluation | Check that each answer is grounded in the supplied document and follows the requested structure. |
Anthropic recommends setting success criteria, an empirical testing method, and a first-draft prompt before beginning prompt engineering. Criteria should be observable enough that two reviewers—or an automated check and a reviewer—can identify whether a response passed.
3. Write an explicit prompt
Specify the operation, audience, inputs, constraints, and output format. Avoid relying on implied intent. If a request has multiple requirements, list them separately so you can inspect whether the response satisfies each one.
You are a technical support assistant.
Task:
Answer the user's question using only the product notes below.
Requirements:
- Give the direct answer first.
- Include numbered steps when the user must perform actions.
- If the notes do not contain the answer, say what is unknown.
- Do not invent product features, URLs, or outcomes.
Product notes:
<notes>
{{product_notes}}
</notes>
User question:
<question>
{{question}}
</question>
Return:
A concise answer, followed by steps if needed.
This example makes the role, source of truth, handling of missing information, and output structure explicit. Adapt the boundaries and wording to your model and application; formatting conventions are aids to clarity, not guarantees of compliance.
4. Give the model the context it needs
Include the task-specific facts, documents, code, definitions, or constraints that the model cannot safely infer. Keep instructions distinct from supplied content. Headings, lists, Markdown, and XML-style tags can help signal where instructions end and source material begins, especially in a long prompt.
- Provide the actual inputs the model must use, not just a description of them.
- Identify which source is authoritative when inputs may conflict.
- State how to handle missing, ambiguous, or conflicting information.
- Keep unrelated background out; extra context can obscure the task.
- For dynamic values, use clearly named placeholders and validate values in application code.
Context is not a substitute for application-side controls. If an output must follow a strict schema, validate it after generation and handle invalid responses. A prompt can request a format; the application remains responsible for deciding whether the returned data is usable.
5. Use examples when they clarify the target
Few-shot examples demonstrate the response pattern you want, including format, scope, tone, and how to handle edge cases. Choose examples that resemble real inputs and use consistent formatting. Include examples for important failure cases when those behaviors matter.
Example input:
Question: Can I export a report as a PDF?
Example output:
Answer: Yes. Open the report and choose Export, then PDF.
Example input:
Question: Can I export a report as a spreadsheet?
Example output:
Answer: The supplied product notes do not say whether spreadsheet export is available.
More examples do not automatically improve a prompt. A model may overfit to a repeated pattern or an overly narrow set of examples. Test whether examples improve performance on cases beyond the examples themselves, and remove examples that add length without clarifying the target.
6. Evaluate and revise systematically
- Build a representative case set. Include ordinary requests, ambiguous inputs, boundary cases, and known failure cases.
- Run a baseline. Save the prompt, model identifier or snapshot, relevant settings, inputs, and outputs.
- Score against criteria. Record pass/fail or a rubric for correctness, completeness, format, and other task-specific requirements.
- Classify each failure. For example: missing context, ambiguous instruction, ignored constraint, incorrect reasoning, invalid format, or capability limitation.
- Change one meaningful thing at a time where practical. This makes it easier to tell whether an edit helped.
- Re-run the same cases and check for regressions. Add newly discovered failures to the evaluation set.
Automated checks can verify mechanical requirements such as valid JSON, required fields, or forbidden strings. Human review may still be needed for correctness, usefulness, and nuanced quality. OpenAI recommends using test and evaluation suites to monitor behavior as prompts or models change; Anthropic likewise emphasizes empirical evaluation against established criteria.
7. Choose the right fix when a prompt fails
| Observed failure | First thing to inspect | Possible next step |
|---|---|---|
| The response omits a required detail | Was the detail in the input, and is it clearly required? | Supply the missing context or make the requirement explicit; add a test case. |
| The response uses the wrong format | Is the requested format concrete and unambiguous? | Show a small example, use structured output support if available, and validate in code. |
| The response invents facts | Does the prompt identify trusted sources and define what to do when information is absent? | Provide source material, require uncertainty to be stated, and check claims against sources. |
| Results vary too much | Are model snapshots, settings, and test inputs held consistent? | Evaluate on the production model and consider pinning a snapshot when consistent behavior matters. |
| Quality remains poor despite clear instructions | Does the model have the capability and context the task requires? | Compare models or redesign the application; more prompt text may not solve a capability mismatch. |
| Latency or cost is too high | Is the chosen model or application path more capable than the task needs? | Evaluate a faster or lower-cost model, reduce unnecessary context, or change the workflow while checking quality. |
Prompt engineering is only one tool. Anthropic notes that not every failed evaluation is best fixed through prompt changes; model selection may improve latency or cost more directly. Compare options on your task’s success criteria, instruction needs, version stability, latency, cost, context handling, and output reliability. There is no universal model ranking supported by the available sources.
8. Maintain prompts as application code
Once a prompt affects production behavior, version and review it like other application logic. Store it in source control, keep representative fixtures, and run evaluations when changing the prompt or model. Use typed inputs or schemas for dynamic values, validate outputs, and deploy through the normal release process.
- Record the prompt version and model snapshot used for each evaluation.
- Test prompt changes against the same representative cases before rollout.
- Keep input construction and output validation in application code.
- Monitor real failures and add safe, representative examples to the evaluation set.
- Re-check current provider documentation because APIs and model workflows change.
9. A compact prompt review checklist
- Is the task stated as a clear action?
- Have success and failure criteria been written down?
- Does the prompt include the necessary context and identify trusted sources?
- Are constraints and output format explicit?
- Would an example remove ambiguity, and has its value been tested?
- Has the prompt been evaluated on representative and edge cases?
- Is the deployed model version the one that was evaluated?
- Could a model or application change solve the problem better?
10. Or skip the browser setup
When prompt evaluation depends on capturing rendered pages—for example, to inspect visual output across URLs—browser setup and capture code become another piece to maintain. ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));
Cookie and consent banners are accepted like a visitor and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. The MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently asked questions
Is prompt engineering just prompt wording?
No. It includes defining success, supplying context and examples, testing outputs, and deciding whether the prompt is the right place to solve a failure.
Will one prompt work with every model?
Not necessarily. Model types and versions can respond differently, so validate prompts on the model and version used in your application.
Should I always add examples?
No. Add examples when they make the expected pattern clearer, then measure whether they help on representative cases.
When should I stop editing the prompt?
When failures point to missing model capability, unsuitable latency or cost, or an application design issue, evaluate those changes instead of extending the prompt indefinitely.


