My First Look at the OpenAI API and Prompt Engineering
I've used AI models from the outside for quite some time. I've asked them questions, given them instructions, generated code with them, and generally treated them as tools.
But using an AI model through an application is a different experience.
Instead of opening a chat interface and typing something into a box, I can write code that sends an input to a model, receives its response, and then does something with that response.
That sounds straightforward.
And at the most basic level, it is.
But once I started looking into how this actually works, I realized that there is quite a bit more going on than simply:
ask the model a question
↓
get an answer
There are models to choose from, different ways of providing instructions, different message roles, different ways of structuring context, prompt caching, few-shot examples, context windows, reasoning models, evaluations, versioning, and a whole set of other ideas that become important once the model becomes part of an actual application.
So this is my first attempt at understanding the OpenAI API and, more specifically, the ideas around prompt engineering.
I'm writing this from the perspective of someone who is still learning it.
There are things here that I understand, things that I only partially understand, and things that I am deliberately learning as I go.
So what is the OpenAI API?
At a high level, the idea is surprisingly simple.
OpenAI provides models that applications can interact with through an API.
Instead of manually opening a chat interface, my application can send an input to a model and receive a response programmatically.
The basic idea looks something like this:
My Application
↓
Send input + instructions
↓
OpenAI API
↓
Model processes the request
↓
API returns a response
↓
My application uses the result
This means the model can become another component inside a software system.
For example, an application could use a model to:
- answer questions
- summarize documents
- generate or transform text
- write or explain code
- extract information
- classify content
- generate structured data
- reason about a problem
- interact with tools
- or perform other tasks that benefit from language-model capabilities
The important part for me is that the model isn't necessarily the entire application.
It can simply be one component inside the application.
That is already a different way of thinking about AI from simply opening a chatbot and asking it questions.
My first API call
A very simple Python example looks something like this:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.5",
input="What are LLMs?",
)
print(response.output_text)
The exact model name available to me can change over time, so the important thing to understand here isn't the particular model name.
The important part is the flow.
- Create an OpenAI client.
- Call the Responses API.
- Choose a model.
- Provide some input.
- Receive a response.
- Read the generated text from
response.output_text.
The response itself isn't simply a string.
The API returns a structured response object. The SDK provides convenient properties for accessing commonly needed pieces of that response.
This distinction becomes important once I start building a real application.
I might not only care about the final text.
I might also care about things such as:
- what the model returned
- what type of output it produced
- whether it requested a tool
- metadata associated with the response
- usage information
- errors
- or other structured information
So even this tiny example is already teaching me something:
The model interaction is structured. The text I see is only one part of the response.
Models are not all the same
One of the first things I noticed is that I don't simply have one universal model that I use for everything.
There are different models designed with different capabilities, performance characteristics, and tradeoffs.
In broad terms, I can encounter models optimized for things such as:
- general-purpose generation
- speed
- lower latency
- lower cost
- difficult reasoning tasks
- coding
- multimodal tasks
- or other specialized capabilities
So choosing a model becomes an engineering decision.
The question isn't simply:
"Which model is the best?"
A better question is:
"Which model is appropriate for this particular task and its requirements?"
A simple classification task might not need the same model as a difficult reasoning problem.
A latency-sensitive application might make a different choice from an application where accuracy is much more important than response time.
An application processing millions of requests might care much more about cost than a small internal tool.
This means model selection is part of application design rather than something I should treat as an afterthought.
Then there is prompt engineering
Once I understood that I could call a model from code, the next question was:
How do I actually tell the model what I want?
This is where prompt engineering comes in.
My current understanding is that prompt engineering is the process of designing the instructions and context given to a model so that it is more likely to produce useful and consistent results.
That sounds simple, but it becomes surprisingly deep.
A prompt isn't necessarily just a question.
It can contain:
- instructions
- context
- examples
- constraints
- desired output formats
- data that the model should operate on
- information about the role or task the model is supposed to perform
So prompt engineering starts looking less like:
"How do I ask ChatGPT a good question?"
and more like:
"How do I design a good interface between my application and a language model?"
That distinction feels important.
Models aren't deterministic functions
One reason prompt engineering matters is that model outputs aren't simply deterministic functions in the same sense as a normal function in a programming language.
If I write:
def add(a, b):
return a + b
then:
add(2, 3)
gives me the same result under the same conditions.
Language models don't work quite like that.
Their generated outputs can vary.
Even when the input is similar, the response can potentially be different.
This means that writing a prompt isn't only about getting one good response.
If I'm building an application, I care about whether the model behaves well across many different inputs.
That leads naturally to another concept:
evaluation.
A rough development loop might look like this:
Write instructions
↓
Run representative inputs
↓
Evaluate outputs
↓
Find failures
↓
Improve the prompt/system
↓
Test again
↓
Repeat
This is why I think prompt engineering is partly an engineering discipline rather than just an art of finding clever wording.
There is experimentation involved, but there should also be testing and measurement.
Instructions are another way to guide the model
One thing that confused me initially was the difference between simply putting instructions into the input and using the separate instructions parameter.
For example:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.5",
instructions="Talk like a cowboy.",
input="Explain the significance of Texas in American history.",
)
print(response.output_text)
The idea is that instructions can provide higher-level guidance about how the model should behave, while input supplies the actual task or content being processed.
This separation starts to make sense when I stop thinking about a prompt as one giant string.
Instead, I can think of the model interaction as having different kinds of information with different purposes.
For example:
- What is the model supposed to be?
- What should it do?
- What should it avoid?
- What information does it have?
- What does the user want?
- What should the output look like?
Those are different questions.
A well-designed application can represent those differences instead of putting everything into one enormous paragraph.
Message roles
Another concept I initially encountered without really understanding was message roles.
In a conversation-style input, messages can have roles that communicate where the message came from and what purpose it serves.
For example:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.5",
input=[
{
"role": "developer",
"content": "Talk like a pirate."
},
{
"role": "user",
"content": "Are semicolons optional in JavaScript?"
}
],
)
print(response.output_text)
Here, there are two different messages.
The developer message provides application-level instructions.
The user message contains the task being requested by the person using the application.
This isn't just metadata attached to the messages.
The roles are part of how the model interprets the interaction.
At a simplified level, I can think about the relationship like this:
Application developer
↓
Defines application behavior
↓
Developer instructions
↓
Establish application-level guidance
↓
User
↓
Provides the actual task
↓
Model
↓
Produces a response within the applicable instructions
There is a hierarchy to these instructions, so the model isn't supposed to treat every piece of text as having exactly the same authority.
This was one of those concepts that I had probably encountered before without actually stopping to ask:
"Why does the role exist in the first place?"
Now I'm starting to see it as part of the structure of the model interaction rather than just another field in a JSON object.
Structuring a prompt
Another thing that surprised me is how much structure can be useful inside a prompt.
Instead of writing one enormous block of text, I can organize the prompt into distinct sections.
A developer instruction might conceptually look like this:
# Identity
You are a coding assistant that helps developers write
clear Python code.
# Instructions
- Prefer readable solutions.
- Explain important tradeoffs.
- Do not invent APIs.
# Examples
<user_query>
How should I parse this file?
</user_query>
<assistant_response>
...
</assistant_response>
# Context
The application runs on Python 3.13.
The exact format isn't magic.
Markdown headings can make the structure clear to both me and the model.
XML-style tags can clearly indicate where one piece of content starts and ends.
For example:
Summarize the document below.
<document>
...
</document>
The important idea is not:
"XML is magic."
The important idea is:
"Make the structure and boundaries of the information clear."
This becomes especially useful when a prompt contains many different pieces of information.
Identity, instructions, examples, and context
One useful structure I've come across is to think about a developer message in four broad sections.
Identity
This describes what the model is supposed to be doing.
For example:
You are a coding assistant that helps developers write Python code.
This establishes the role and general purpose.
Instructions
These describe how the model should perform the task.
For example:
- Prefer readable code.
- Explain important tradeoffs.
- Don't invent APIs.
- Return only Python code when code is requested.
The instructions tell the model what behavior is expected.
Examples
Examples show the model what a desired interaction or output looks like.
Instead of only saying:
"Return the answer in this format."
I can demonstrate the format with an example.
For example:
Input:
Convert this sentence into a formal email.
Output:
Dear...
The example provides something concrete for the model to follow.
Context
Context is additional information that the model needs to perform the task.
This might include:
- company-specific information
- documents
- database records
- current application state
- user information
- product information
- internal documentation
- or other information that isn't something I expect the model to already know
Thinking in these sections makes prompts much easier for me to reason about.
Instead of thinking:
"This is my prompt."
I can start thinking:
Identity
+
Instructions
+
Examples
+
Context
+
User input
↓
Model response
That is a much more useful mental model.
Markdown and XML are not the model's programming language
One thing I want to be careful about is not turning a useful technique into a superstition.
Markdown and XML-style delimiters don't magically make a model "understand" a prompt in some completely different way.
Their main benefit is that they make the structure of the input clearer.
For example, instead of writing:
Summarize this document and don't mention confidential information
document starts here ...
I can make the boundary explicit:
Summarize the document below.
<document>
...
</document>
Now the distinction between the instruction and the supplied data is much clearer.
The model still processes tokens.
It doesn't suddenly become an XML parser simply because I put <document> around something.
The value is that the structure gives the model useful patterns and makes the prompt less ambiguous.
And it makes the prompt much easier for me to read and maintain too.
That last part is important.
A prompt is not only written for the model.
If I'm going to maintain an AI-powered application, other developers and future me also need to understand it.
Few-shot learning
Another concept that caught my attention is few-shot learning.
My initial understanding was:
Instead of fine-tuning the entire model, I can provide examples of inputs and desired outputs inside the prompt, and the model can use those examples as a pattern for the current task.
For example, suppose I want a model to classify support tickets.
I could provide examples like:
Input:
"My payment was charged twice."
Output:
billing
Input:
"I cannot reset my password."
Output:
account
Input:
"The application crashes when I upload a PDF."
Output:
technical
Then I could provide a new input:
Input:
"My credit card was charged for the same order two times."
The examples provide a pattern that the model can use when generating the next answer.
This is different from fine-tuning the model's underlying parameters.
The examples are part of the input context for that particular request.
This is why the term few-shot makes sense.
I'm giving the model a few examples of what I want rather than changing the model itself.
There are also zero-shot and one-shot approaches.
Very roughly:
Zero-shot
↓
Give the task without examples
One-shot
↓
Give one example
Few-shot
↓
Give several examples
The examples themselves become part of the prompt.
That means they are effectively part of the specification I'm giving the model.
So I should choose them carefully.
More examples aren't automatically better
My first instinct was:
"If examples help, then I should give the model as many examples as possible."
But that isn't necessarily true.
More examples mean more input tokens and more context for the model to process.
What matters is whether the examples are useful and representative.
I'd rather have a smaller collection of diverse, high-quality examples than a huge collection of repetitive examples that don't add much information.
For example, if I am trying to teach a model how my application classifies support tickets, I probably want examples covering different kinds of billing, technical, and account problems.
I don't necessarily need fifty examples that all say:
"I was charged twice."
So the question becomes:
Which examples teach the model the pattern I actually care about?
That is a much more useful question than simply asking how many examples I should provide.
Context windows
All of this eventually leads to another important concept:
How much information can I actually give the model at once?
This is where the context window comes in.
Models process information in units called tokens.
A token isn't exactly the same thing as a word.
Depending on the text, a token can represent part of a word, a whole word, punctuation, or other pieces of text.
The context window represents the amount of input and conversational context that a model can work with for a request, subject to the limits of the particular model and API configuration.
This matters because everything I put into the request becomes part of the information the model has to process.
If I keep adding:
- instructions
- examples
- documents
- conversation history
- user information
- tool results
- retrieved information
the context can become very large.
So prompt engineering is partly about deciding:
What information does the model actually need?
Not:
"What information can I possibly throw at the model?"
That distinction matters.
Prompt caching
This leads naturally to prompt caching.
Imagine an application where the same large set of instructions and examples is sent repeatedly, while only the user's actual question changes.
For example:
Large stable instructions
+
Large collection of examples
+
User's new question
Then another request:
Large stable instructions
+
Large collection of examples
+
Another user's question
And another:
Large stable instructions
+
Large collection of examples
+
Another question
Sending the same large prefix repeatedly can be inefficient.
Prompt caching can allow supported repeated input content to be reused, potentially reducing latency and input costs.
The conceptual idea is:
Stable prompt prefix
+
Dynamic input
↓
API request
Stable prompt prefix
+
Different dynamic input
↓
Another API request
Cached stable content
↓
Potentially reduce repeated processing/cost
This means prompt structure can become an engineering concern, not just a readability concern.
If I have a large static instruction set that I send repeatedly, I want that stable content to be positioned consistently so that it can benefit from caching.
Something that initially looked like a prompting technique starts connecting to:
- performance
- latency
- cost
- application architecture
And that is exactly the kind of connection I want to understand.
Prompts as code
Another idea that I'm learning about is treating prompts as part of the application itself.
This makes intuitive sense to me.
If the behavior of my application depends on a large instruction such as:
Summarize customer feedback.
Focus on:
- recurring problems
- positive feedback
- unusual complaints
Return the result as bullet points.
then that instruction is effectively part of my application's behavior.
It makes sense that I should be able to:
- version it
- review changes to it
- test it
- deploy changes to it
- roll changes back
- evaluate whether a new version is actually better
That starts looking a lot like normal software development.
The development process could look like:
Write prompt
↓
Commit to Git
↓
Code review
↓
Run tests/evals
↓
Staging
↓
Production
↓
Monitor
↓
Improve
This is one of the places where AI application development starts to feel much more like ordinary software engineering.
A prompt isn't necessarily just some text I wrote in a chat box.
In a production application, it can be an important part of the application's behavior.
Versioning prompts
If I change:
Summarize customer feedback.
into:
Summarize customer feedback in concise bullet points,
grouping recurring issues together.
I've changed application behavior.
That change should be visible in version control just like a normal code change.
For example, Git could show something as simple as:
- Summarize customer feedback.
+ Summarize customer feedback in concise bullet points,
+ grouping recurring issues together.
That's useful because someone reviewing the code can immediately see that the behavior of the AI component changed.
This also means that prompt changes should be reviewed rather than casually edited without knowing what changed.
A prompt can affect thousands of model interactions just as a change to a function can affect thousands of requests.
That makes versioning and testing especially important.
Testing prompts
This is probably one of the biggest differences between writing a normal deterministic function and writing an AI-powered component.
I can't simply test:
assert summarize("hello") == "expected exact string"
and assume that I've completely tested the system.
The model can generate different valid outputs.
So I need to think more carefully about what "correct" means.
For example:
Input:
Customer says their payment was charged twice.
Expected behavior:
- identify this as a billing issue
- acknowledge the duplicate charge
- respond professionally
- do not invent a refund policy
Now I can evaluate whether the response satisfies those requirements rather than checking whether the response is character-for-character identical to an expected string.
This is where evaluations, or evals, become important.
A production AI system needs a way to answer:
"Did this change actually make the model's behavior better?"
That is a much harder question than:
"Did my function return the expected string?"
And that is one of the areas I want to explore much more deeply.
Prompt changes become software changes
I'm starting to see a useful pattern here.
A prompt can go through a normal development lifecycle:
Development
↓
Testing
↓
Evaluation
↓
Staging
↓
Production
↓
Monitoring
↓
Improvement
And just like normal software, I could eventually introduce things such as:
- feature flags
- A/B testing
- automated evaluations
- regression tests
- versioned releases
- monitoring
- gradual rollouts
That is a much more useful way for me to think about prompt engineering than:
"I need to find the perfect prompt."
There probably isn't one perfect prompt.
There is a prompt that performs well for a particular task, model, input distribution, and set of requirements.
And I need a way to know whether it continues to perform well when something changes.
Maybe the model changes.
Maybe my application changes.
Maybe the user inputs change.
Maybe the prompt changes.
Maybe the requirements change.
Again, evaluations become important.
Reasoning models vs GPT models
Another distinction that I am still learning is the difference between reasoning-oriented models and GPT-style general-purpose models.
One mental model that helped me understand the distinction is this:
A reasoning model can be thought of more like a senior coworker: give it the goal and important constraints, and allow it to work out more of the details.
And:
A general-purpose GPT model can be thought of more like a coworker who benefits from more explicit instructions about exactly what the desired output should look like.
This is a simplification.
It isn't a rule that says:
"Reasoning models need vague prompts."
or:
"GPT models need extremely detailed prompts."
The reality is more nuanced.
But the analogy gives me a useful starting mental model.
It also tells me something important:
I shouldn't assume that a prompting technique that works well for one model family will work identically for another.
Different models can have different strengths and different prompting characteristics.
So once again, testing matters.
What about model versions?
There is another thing I hadn't really thought about before:
A model name isn't necessarily enough to describe the exact behavior I'm depending on.
Models can be updated, replaced, or deprecated.
For production applications, this means I need to pay attention to model versions and changes rather than assuming that an AI model is a permanently fixed dependency.
If an application depends heavily on a particular model's behavior, changing the model can be similar to changing an important software dependency.
The right response isn't necessarily:
"Never upgrade."
It's:
"Have I tested what happens when I upgrade?"
That is a much healthier software engineering mindset.
I don't have to be afraid of upgrading.
I just need to understand that an AI model is a dependency and that changing an important dependency can change application behavior.
Again, evaluations become important.
Context is more important than I initially thought
Another lesson I'm beginning to understand is that a model's output depends heavily on the information I give it.
If I ask:
Summarize this.
without actually providing anything to summarize, the model has very little useful context.
But if I provide:
Summarize the following customer feedback.
Focus on:
1. recurring complaints
2. positive feedback
3. requested features
<customer_feedback>
...
</customer_feedback>
I've given the model a much clearer task.
This is why additional context can be so powerful.
The model might have broad knowledge from training, but my application may have information that isn't available from that general knowledge.
For example:
- internal company policies
- private customer information
- current database records
- recent application state
- internal documentation
- product information
- information that wasn't available in the model's training data
Supplying relevant context is therefore a major part of building useful AI applications.
And this starts leading toward other concepts I haven't explored properly yet.
For example:
- retrieval
- embeddings
- vector databases
- retrieval-augmented generation
- tool calling
- external APIs
- structured outputs
I don't want to jump into all of those immediately.
But I can already see why they exist.
I'm starting to see prompts differently
Before learning about all of this, I mostly thought of a prompt as:
"The question I type into an AI."
I'm starting to see it differently now.
In an application, a prompt can be much closer to a structured interface between my software and a model.
A simplified picture in my head now looks something like:
Application behavior
↓
Instructions
↓
Examples
↓
Relevant context
↓
User input
↓
Model
↓
Structured response
↓
Application behavior
And that interface needs engineering around it.
I need to think about:
- what instructions I provide
- how I structure them
- what examples I provide
- what context is relevant
- how much context I can provide
- which model I use
- how I test the output
- how I handle model changes
- how I measure whether the result is actually good
This is starting to feel much more like software engineering.
What I understand so far
I started this topic thinking that using an LLM through an API would mostly be:
send prompt
↓
receive response
That mental model isn't wrong.
It's just incomplete.
A better mental model for me now looks something like:
Application
↓
Choose an appropriate model
↓
Construct instructions and input
↓
Provide relevant context and examples
↓
Send request through the API
↓
Receive structured response
↓
Evaluate and use the output
↓
Monitor behavior
↓
Improve the system
↓
Repeat
And I suspect this is still only the beginning.
There is a lot more to learn
I still have quite a few questions.
How exactly are the different message roles represented internally?
What exactly happens to my input before it reaches the model?
How do tokens actually work?
How does tokenization work for code?
How does the context window interact with long conversations?
How should I design evaluations for non-deterministic outputs?
How should I decide between different models?
When should I use few-shot examples?
When should I retrieve external context instead?
How should prompts be structured in a large production application?
How do tool calls fit into the same picture?
How do structured outputs fit into the model's response?
How does streaming work?
What actually happens when an API request travels from my application to the model and back?
And eventually:
How do all of these pieces come together to build an actual AI-powered application?
I don't have all of those answers yet.
That's fine.
This is the beginning of the learning process.
What I want to take away from this
The biggest thing I want to avoid is treating LLMs as magic.
I don't want my understanding of AI development to stop at:
"Put a good prompt into ChatGPT and hope it works."
I want to understand the engineering underneath that interaction.
A model is a component.
A prompt is part of the interface to that component.
Context provides information the model needs.
Examples can demonstrate desired behavior.
Model selection introduces tradeoffs.
Evaluations tell me whether changes actually helped.
Version control lets me treat prompts and surrounding logic as software artifacts.
And the API gives my application a programmatic way to put all of these pieces together.
I'm still learning the details, but the system is starting to look less like magic and more like software engineering.
And honestly, that's probably the part that interests me the most.
Where I go from here
This was my first pass at understanding the OpenAI API and prompt engineering.
I started with what looked like a very simple idea:
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-5.5",
input="What are LLMs?",
)
print(response.output_text)
But that tiny piece of code opens the door to a much larger system.
There are models.
There are instructions.
There are message roles.
There is context.
There are examples.
There are tokens.
There is caching.
There are evaluations.
There are model versions.
There are APIs.
There are tools.
There are structured outputs.
And there are many more things I haven't understood yet.
So I don't think the next step is to memorize all of these terms.
It's to start taking them apart one by one.
To understand what each piece actually does.
To experiment with it.
To break my own assumptions.
And then to connect it back to the bigger picture.
One API call at a time.