Prompt Engineering for Developers

Prompt engineering is mostly specification writing. Here is a structure that produces reliable output, and the failure patterns that break it.

Most advice about prompt engineering is folklore: be polite, assign a role, offer a tip. Some of it helps by accident. What actually makes prompts reliable is much less glamorous and much more transferable — it is specification writing. The prompts that work consistently are the ones that would also work as a clear brief for a competent human contractor who cannot ask you questions.

This guide gives you a structure that generalises, shows what it changes concretely, and covers how to evaluate improvements without fooling yourself.

Reframe: it is specification, not persuasion

The intuition that you need to coax or flatter the model is a category error. A language model is a conditional text generator: given a context, it continues in the way its training makes most probable. Your prompt is not a request you are making of an agent with feelings; it is the context that shapes the distribution it samples from.

That reframe has immediate practical consequences:

  • Politeness is neutral. "Please" neither helps nor hurts on any measurable benchmark. Write naturally and spend your tokens on content.
  • Role assignment is weak on its own. "You are a senior engineer" does little by itself. What works is the criteria that role implies: what to prioritise, what to reject, what standard to apply.
  • Constraints beat encouragement. "Be thorough" is unmeasurable. "Return at most five findings, each with a line number and a concrete fix" is a specification the model can satisfy and you can verify.
  • What you omit matters as much as what you include. An underspecified prompt leaves the model to guess your intent from its training distribution — which is usually the average of everything, not what you wanted.

The five-part structure

Not every prompt needs all five parts, but knowing which one you have omitted tells you exactly why your output is unreliable.

1. Role and objective

Not "act as an expert" but the specific stance and goal: what perspective to take, and what the output is for. This is cheap and it reliably shifts tone and depth.

2. Context

Everything the model cannot know: your domain, your constraints, your audience, the surrounding situation. This is where most prompts are thinnest, and it is usually the highest-value addition.

3. Input

The material to operate on, clearly delimited. Use XML tags, triple backticks, or unambiguous section headers. Delimiting is not cosmetic — it prevents the model from confusing your instructions with your data, which is also the primary defence against prompt injection when the input comes from users.

4. Constraints

The measurable boundaries: length, format, what to include, what to exclude, what to do when information is missing. Every constraint you can state, state.

5. Output format

State the shape explicitly, and show an example if the format is non-obvious. A schema beats a description. If you need machine-parseable output, say so and specify the schema.

Before and after

Before

Review this code and tell me if there are any problems.

def process(items):
    result = []
    for item in items:
        if item.status == 'active':
            result.append(item.value * 1.2)
    return result

This prompt is underspecified in at least five ways: no language or version, no standard to apply, no severity threshold, no output format, and no instruction about what to do if the code is fine. The model will produce something plausible and generic.

After

You are reviewing a Python 3.12 pull request for a payments service.
Prioritise correctness and security over style.

<code>
def process(items):
    result = []
    for item in items:
        if item.status == 'active':
            result.append(item.value * 1.2)
    return result
</code>

Review for: correctness bugs, unhandled edge cases, and security issues.
Ignore formatting and naming preferences unless they obscure meaning.

Return at most 5 findings. For each: severity (high/medium/low), the line
number, what is wrong, and a specific fix.

If the code has no issues worth reporting, respond with exactly:
"NO_ISSUES_FOUND"

Format each finding as:
[SEVERITY] line N: <problem> → <fix>

What changed concretely: the standard is stated (correctness and security over style); the scope is bounded (three categories, ignore the rest); the output is specified (format, count cap, severity); and there is an explicit negative case — the "say nothing is wrong" instruction, which is the single most reliable way to stop a model from inventing problems to appear helpful.

Controlling output format

If anything downstream parses the output, format is not cosmetic — it is correctness. Four techniques, in order of reliability:

Show, don't describe

One concrete example of the exact output you want beats three paragraphs describing it. Few-shot examples are the most reliable formatting control available, and they also teach tone and granularity.

Use a schema, not prose

Return a JSON object matching this schema. No prose before or after.

{
  "findings": [
    {
      "severity": "high" | "medium" | "low",
      "line": number,
      "problem": string,
      "fix": string
    }
  ],
  "summary": string
}

Use structured output modes when available

Most providers now support JSON schema-constrained or grammar-constrained decoding, which makes invalid output structurally impossible rather than merely unlikely. If your provider supports it, use it — it eliminates an entire class of production failures, and it costs you nothing in flexibility for most use cases.

Parse defensively anyway

Even with structured output, validate before use. Never pass model output straight into a database query, a shell command, or an HTML template. Treat it as untrusted input, because it is.

Failure patterns

Sycophancy

Models tend to agree with positions stated in the prompt. If you write "this approach seems right, what do you think?", you have biased the answer. Present options neutrally and ask for evaluation against stated criteria instead. For a review you actually want to be adversarial, ask for the strongest case against the proposal, then decide yourself.

Invented specifics

When a prompt implies an answer exists, models supply one. Add an explicit escape hatch every time: "If the information is not present in the provided text, say 'not stated'. Do not guess." This one line prevents a large share of hallucinated detail.

Instruction drift in long contexts

Instructions at the very beginning of a long context are followed less reliably than those at the end. If you have a long document plus instructions, consider repeating key constraints after the content, or putting the most important instruction last.

Prompt injection

When any part of your prompt comes from a user, a document, or a web page, that content can contain instructions — "ignore previous instructions and…". Delimit untrusted content clearly, tell the model that content inside the delimiters is data, never instructions, and keep your actual instructions outside those delimiters. This reduces risk substantially but does not eliminate it; for anything security-sensitive, treat model output as untrusted and enforce policy in code.

How to evaluate a change

The most common way teams fool themselves: change a prompt, run it on two examples, see better output, ship it. Two examples is anecdote, and confirmation bias does the rest.

A minimum viable process:

  1. Build a fixed set of 30–100 test cases before you start iterating. Include the hard ones and the ones where the current prompt fails.
  2. Define success per case mechanically. Exact match, schema validity, a regex, or a small judge model. Avoid "looks better to me".
  3. Change one thing at a time. If you change three things and the score moves, you have learned nothing about which change mattered.
  4. Run the full set every time. Prompt changes that improve one case frequently regress another.
  5. Version your prompts like code. Store them in your repository with the evaluation results attached. Prompts are production configuration.

Non-determinism matters here: at temperature 0 output is usually stable but not guaranteed identical across model updates. If you have a single test case whose result flips between runs, your evaluation is measuring noise.

Anti-patterns

  • Prompt as junk drawer. Prompts accumulate instructions over months and are never pruned. Every instruction costs input tokens on every request, and contradictory instructions degrade output. Audit periodically.
  • Magic phrases. "Take a deep breath", "you will be tipped", "think step by step" — these circulate because they occasionally help on specific benchmarks, not because they are general principles. Test them on your own workload or skip them.
  • Solving in the prompt what belongs in code. If a task needs deterministic transformation, sorting, arithmetic or validation, write code for it and let the model handle the part that genuinely requires language understanding.
  • One giant prompt. A prompt trying to classify, extract, summarise and translate at once will do all four worse than four focused prompts — and will be far harder to debug. Decompose, and use code to chain the steps.
  • No fallback. Production systems need a defined behaviour when the model returns something unusable: retry, degrade, or hand off to a human. Design that path before you need it.

Frequently asked questions

No measurable effect on output quality. Politeness is harmless but consumes tokens; spend them on context and constraints instead.

As long as it needs to be and no longer. Under-specification is far more common than over-specification, but every token costs money on every request, so audit prompts for accumulated instructions periodically.

Yes when output format, tone or granularity matters and is easier to show than describe. Two or three well-chosen examples typically outperform longer prose instructions on formatting.

Give it an explicit escape hatch: 'if the answer is not in the provided text, say so'. Require citations or quotes from the source. And verify anything factual against the source system in code.

Start at 0 for anything deterministic — extraction, classification, structured output. Raise it only for tasks where you genuinely want variation, and then evaluate whether the variation is worth the non-determinism you are introducing.