ScreenshotNeo

BlogAI agents

Free AI Agent Tutorial: Build Your First Agent

Build a first AI agent in Python or JavaScript, then add tools, state and workflows. Learn what free inference really means and how to inspect a run.

By the ScreenshotNeo team29 September 202611 min read

Free AI Agent Tutorial: Build Your First Agent

To build your first AI agent for free, start with one narrow task, give an agent concise instructions, and run it once. This tutorial shows a first run in Python and JavaScript, explains which parts can be free, then walks through tools, state, workflows, tracing, and deployment. You do not need an agent framework for every project: a single model call may be enough until the program needs to choose and use tools.

“Free” has two meanings here. The SDKs are software you can install, while model inference may be free only within a provider’s limited tier, or free of hosted per-call charges when you run a model locally. Limits, model availability, and pricing can change. Check the linked provider pages before building around a particular allowance.

1. Build the smallest working AI agent

The shortest beginner path is Python with the OpenAI Agents SDK. You will create a virtual environment, install the package, put your API key in an environment variable, and run one prompt. The JavaScript quickstart below does the equivalent with npm. Both official quickstarts are designed to get a first agent running in Python or JavaScript. Python quickstart · JavaScript quickstart.

Python: first run

  1. Create and activate an environment, then install the SDK.
  2. Set OPENAI_API_KEY in your shell or environment manager.
  3. Save the code as first_agent.py and run it.
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
# .venv\Scripts\Activate.ps1
pip install openai-agents

# macOS or Linux:
export OPENAI_API_KEY="your-api-key"
# Windows PowerShell:
# $env:OPENAI_API_KEY="your-api-key"

# Save the following as first_agent.py
import asyncio
from agents import Agent, Runner

async def main():
    tutor = Agent(
        name="History tutor",
        instructions=(
            "You are a concise history tutor. Answer the question clearly, "
            "define unfamiliar terms, and say when you are uncertain."
        ),
    )
    result = await Runner.run(tutor, "Why did the printing press matter?")
    print(result.final_output)

if __name__ == "__main__":
    asyncio.run(main())

Keep the key out of the file and source control. If you use a virtual environment, install the SDK inside the activated environment so the agents import resolves from the same Python installation that runs the script.

JavaScript: equivalent first run

The JavaScript SDK uses npm and Zod for schemas when you later define typed tools. This first example needs only the agent package.

mkdir first-agent
cd first-agent
npm init -y
npm install @openai/agents zod

# macOS or Linux:
export OPENAI_API_KEY="your-api-key"
# Windows PowerShell:
# $env:OPENAI_API_KEY="your-api-key"

# Save as index.mjs
import { Agent, run } from "@openai/agents";

const tutor = new Agent({
  name: "History tutor",
  instructions:
    "You are a concise history tutor. Answer clearly, define unfamiliar terms, and say when uncertain.",
});

const result = await run(tutor, "Why did the printing press matter?");
console.log(result.finalOutput);

# Run:
node index.mjs

Use .mjs for this example so Node treats it as an ES module. If you prefer .js, set "type": "module" in package.json. Store keys in your deployment’s secret manager in production; do not expose a server key in browser JavaScript.

2. What makes this an agent?

At this stage, the agent has three ingredients: instructions that set its role, a model that generates the response, and a runner that executes the turn and returns output. In the Python SDK, Runner.run produces a result with final_output; the JavaScript quickstart exposes the equivalent finalOutput. The runner and result also give you run history and tracing hooks as the application grows. See the Agents SDK documentation.

A first agent starts with instructions and one run; a tool adds a controlled action when the task needs it.
A first agent starts with instructions and one run; a tool adds a controlled action when the task needs it.

A plain model call can answer a question, but an agent becomes useful when your program can let it select an action, receive the action’s result, and continue toward a goal. Add those capabilities one at a time. A role prompt alone does not guarantee factual accuracy, safe behavior, or access to current information.

3. Add one tool when the task requires it

A tool lets the model request a specific operation your code controls. A good first tool has a narrow purpose, typed inputs, predictable output, and an understandable failure result. For example, an FAQ agent might call a function that looks up a product’s published return window. Keep authoritative policy data in your application rather than asking the model to invent it.

For a function tool, define a name and description the model can choose, validate the input schema, run the function, and return a result. Handle expected errors in the function and avoid returning secrets or unnecessary private data. The SDK supports function tools and hosted tools; see its tools guide for current syntax and options.

  1. Describe the decision. Write down when the agent should call the tool and when it should answer without one.
  2. Constrain inputs. Use an enum or schema where practical, set length limits, and reject invalid identifiers.
  3. Make side effects deliberate. For actions such as sending a message or changing a record, require application authorization and consider a user confirmation step.
  4. Return useful errors. Distinguish “not found” from a temporary service failure; do not silently fabricate a successful result.
  5. Test the boundary. Try malformed arguments, no match, duplicate requests, and a downstream timeout.

Some tasks do not need an agent to choose tools. If the operation is fixed and deterministic, call the function directly from your application and use a model only for the language task. This is often simpler to debug and cheaper to operate.

4. Add conversation state and memory deliberately

The first example is stateless: it runs one prompt and ends. A conversational application must decide what context to carry into the next turn. Session history is a record of prior interaction; memory is selected information retained for later use. They are related but not identical.

  • Short conversation: pass or store the recent messages needed to answer follow-up questions.
  • Long conversation: summarize older turns or retrieve relevant records instead of sending the entire transcript every time.
  • Cross-session preference: store only useful, consented facts with a clear retention policy and a way to correct or delete them.
  • Workflow state: persist explicit task status in your application database when a job must survive a process restart.

Do not treat model context as durable storage. Define what happens when state is missing, stale, too large, or inconsistent. Microsoft’s staged tutorial is a useful learning sequence: first agent, tools, conversations, memory, workflows, harness, and hosting. Microsoft Agent Framework tutorial.

5. Use handoffs and workflows only when needed

A single agent is easier to understand than a group. Add multiple specialists when they have distinct responsibilities, or use a workflow when the process has explicit stages, branches, or checks. Handoffs let one agent route work to another; agents-as-tools allow a coordinating agent to invoke a specialist for a bounded subtask. Guardrails and structured outputs can help enforce boundaries and make results easier for software to consume. The SDK documents handoffs, agents-as-tools, guardrails, and structured outputs.

Before splitting an agent, write the steps in ordinary code or a diagram. If a normal function, queue, or state machine is clearer, use it. A workflow that can send emails, spend money, update customer records, or access private systems needs authorization checks outside the model as well as clear logging and recovery paths.

6. Choose a free model path

There is no universal unlimited-free option. The practical choices are a capped hosted free tier or local inference on hardware you control.

Path Useful for Tradeoffs to check
OpenAI Agents SDK with a hosted model Learning the agent runner, tools, and tracing Model API usage may be billed; check current account access and OpenAI API pricing.
Gemini API free tier Prototyping eligible models with an allowance Free-tier model eligibility, request limits, and terms vary. Check Gemini API pricing and rate limits.
Local model via Ollama Learning or experiments without a hosted per-call charge Requires suitable memory and compute; speed and output quality depend on hardware and model. Review the Hugging Face local apps guide and the model’s license.
Hugging Face inference providers Trying hosted models through a common interface The documented free-user allowance is small and may change; verify current pricing.

Provider limits and prices change, so treat free access as a way to learn and prototype, not a production capacity guarantee. For a production estimate, record the model, input and output token volumes, retry rate, and tool-call frequency, then calculate against the provider’s live pricing page. Also account for hosting, storage, observability, and any external tools your application calls.

7. Inspect the run before expanding it

Once the first answer works, inspect what happened. Review run history or tracing to identify model calls, tool calls, handoffs, duration, and failures. A trace helps answer practical questions: Did the agent call the expected tool? Did the tool return an error? Did a retry duplicate an action? Did a specialist receive more context than it needed?

For each important task, create a small evaluation set of representative prompts and expected behavior. Include normal requests, ambiguous requests, missing data, adversarial input, and tool failure. Evaluate outcomes after changing instructions, models, schemas, or prompts. A successful demonstration is not evidence that the agent behaves correctly across the cases your users will send.

8. Capture web pages for an agent’s visual workflow

If your agent needs to inspect a webpage visually, a screenshot is a useful tool result: capture a page, then pass the image to a vision-capable model or use it as an artifact for a human review. A browser-based implementation gives you control, but adds browser setup, navigation waits, viewport choices, and cleanup of overlays. The ScreenshotNeo website screenshot API can also be called as a tool from your agent when a web page is part of its task.

A screenshot can become a visual tool result for an agent, with overlays removed before capture.
A screenshot can become a visual tool result for an agent, with overlays removed before capture.

For a do-it-yourself browser route, Playwright can navigate to a page and save a screenshot. Install its browser binaries using the official Playwright installation guide, then use this script:

npm init -y
npm install playwright
npx playwright install chromium

# Save as capture.mjs
import { chromium } from "playwright";

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto("https://example.com", { waitUntil: "networkidle", timeout: 30000 });
  await page.screenshot({ path: "page.png", fullPage: true });
} finally {
  await browser.close();
}

# Run:
node capture.mjs

For frequently updated pages, network idle may never occur because analytics, streaming, or long polling keep requests active. In that case, wait for a meaningful selector or a short, bounded delay, and set an explicit timeout. Full-page screenshots can be large; use a viewport capture when only the visible region is needed. Do not send a screenshot containing personal or confidential information to a model unless the application’s privacy and retention requirements allow it.

9. Or skip the browser setup

ScreenshotNeo turns a URL into a screenshot image or PDF with one GET request. See the API documentation for parameters. For an agent tool, make the request from your server, check the response, and pass the image bytes or a controlled URL to the next step.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer())));

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents, including Claude and Cursor, take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for 1,000 screenshots a month, with no card.

10. Troubleshooting common first-run problems

Symptom Likely cause Fix
Missing API key or authentication error The variable is unset in this shell, misspelled, or unavailable to the process. Set OPENAI_API_KEY in the same terminal, restart the process, and verify the secret name in your environment settings. Never print the key to logs.
ModuleNotFoundError: agents Package installed into a different Python environment. Activate the intended virtual environment, then run python -m pip install openai-agents with that interpreter.
Node says import is unsupported Node is treating the file as CommonJS. Use an .mjs extension or configure "type": "module" in package.json.
Rate limit or quota response Account/model limits or billing configuration prevent the request. Check the provider’s current quota and rate-limit page, reduce concurrency, use bounded backoff for transient limits, and avoid retrying permanent billing errors.
Agent answers without using a tool Tool description is vague, inputs do not fit, or instructions allow a direct answer. Make the tool’s purpose and invocation conditions explicit, then inspect traces and test prompts that require the tool.
Repeated or harmful side effect Retries or model decisions can repeat an operation. Use idempotency keys, authorization checks, confirmation for consequential actions, and application-level limits.
Slow or stuck browser capture Network idle may not be reached, the page is heavy, or navigation is blocked. Set a timeout, wait for a specific selector or bounded delay, and inspect navigation errors. Close the browser in a finally block.

11. A practical checklist before sharing your agent

  • It solves one defined task and has a concise instruction set.
  • Credentials are stored outside code and logs.
  • Tools validate inputs and return explicit errors.
  • Side effects are authorized, bounded, and safe to retry.
  • Conversation state has a size, retention, and deletion policy.
  • Representative success and failure cases are evaluated.
  • Provider limits, model terms, and production costs are checked against current documentation.
  • Timeouts, cancellation, retries, and a useful user-facing failure path are in place.

12. Frequently asked questions

Can I build an AI agent without paying for an API?

Yes, for learning you can use a provider’s eligible capped free tier or run a local model. Neither path guarantees unlimited capacity; the local option shifts the cost to hardware and setup.

Should a beginner start with Python or JavaScript?

Choose the language you can already run and debug. Python has a short virtual-environment workflow; JavaScript fits npm projects and web applications. The agent concepts in the two examples are the same.

Do I need an agent framework?

No. Start with a direct model call if the task is one request and response. A framework becomes helpful when you need tool routing, handoffs, run history, or structured orchestration.

What should I learn after the first run?

Add one read-only tool, inspect its trace, then decide whether your application needs conversation state. Add memory or workflows only for a specific requirement.

Can an agent take a screenshot?

Yes. Implement a browser capture function or connect an MCP screenshot server, then give the agent a narrow tool description and a controlled way to pass the result onward.