Using AI Agents to Turn Task Descriptions Into Structured Data
Learn how to turn natural-language tasks into validated JSON with schemas, grounded extraction, failure handling, and production checks.

AI agents can turn a natural-language task description into structured data when you give them an explicit record schema, require grounded extraction, and validate the result before your application acts on it. The reliable pattern is: define the fields, ask the agent to extract only supported values, parse the structured response, run application-level checks, and route uncertainty or failures for review.
A valid JSON object proves that the output matches a shape. It does not prove that the agent understood every detail, found every relevant field, or avoided an unsupported inference. Treat schema compliance, factual correctness, and completeness as separate properties.
1. Define the record before you prompt
Start with the data contract your application needs. For every field, decide its name, type, required or optional status, allowed values, and representation of missing information. A schema is more useful when it expresses business rules instead of merely asking for “some JSON.”

| Design question | Example decision |
|---|---|
| What is the field called? | priority |
| What type is it? | String enum |
| Can it be absent? | Optional; use null when unknown |
| What values are allowed? | low, medium, high, urgent |
| What does “unknown” mean? | The description does not support a value |
| What needs an example? | Date formats, identifiers, and ambiguous labels |
Keep extraction fields distinct from derived fields. For example, store the requested due date separately from a calculated SLA deadline. The agent can extract the former; deterministic application code should usually calculate the latter.
Example JSON Schema
{
"type": "object",
"additionalProperties": false,
"properties": {
"summary": { "type": "string" },
"assignee": { "type": ["string", "null"] },
"priority": {
"type": ["string", "null"],
"enum": ["low", "medium", "high", "urgent", null]
},
"due_date": {
"type": ["string", "null"],
"description": "ISO 8601 date, YYYY-MM-DD"
},
"labels": {
"type": "array",
"items": { "type": "string" }
},
"needs_clarification": { "type": "boolean" },
"evidence": {
"type": "array",
"items": { "type": "string" }
}
},
"required": [
"summary", "assignee", "priority", "due_date",
"labels", "needs_clarification", "evidence"
]
}
The evidence field is useful for review. Ask the agent to quote or paraphrase the specific part of the task that supports each important value. Keep evidence short and store the original task beside the result for auditing.
2. Write an extraction instruction that prevents guessing
Your instruction should explain the field meanings, the boundary between extraction and inference, and the response behavior for missing or conflicting details. A practical instruction looks like this:
You extract a task record from the supplied description.
Use only facts stated in the description. Do not infer an assignee,
priority, date, label, or requirement from general knowledge.
Use null when a scalar value is absent. Use an empty array when no
labels are stated. Set needs_clarification to true when wording is
ambiguous or required information conflicts. Include short evidence
for each populated field. Return only the declared schema.
Give the model a few examples for decisions that people commonly phrase inconsistently. For instance, explain whether “by Friday” means the next Friday in the user’s timezone, whether “ASAP” maps to urgent, and whether “the payments team” is a valid assignee or requires a person identifier.
Distinguish absence, ambiguity, and contradiction
- Absent: the description does not mention a due date. Return
null. - Ambiguous: “ship it next week” has no agreed date. Return
nulland setneeds_clarificationto true. - Contradictory: one sentence says “low priority” and another says “production outage.” Preserve the conflict in evidence and send the record for review.
3. Use constrained structured output where available
When the model or agent framework supports schema-constrained output, pass the schema through that mechanism rather than asking for JSON in ordinary text. The OpenAI Agents SDK documents output schemas that validate and parse agent results. OpenAI’s function-calling documentation describes strict Structured Outputs that match generated arguments to a supplied JSON Schema, and its API guidance covers structured extraction from unstructured inputs. Agents SDK output schemas, function calling and Structured Outputs, and the Structured Outputs guide are primary references.
Google documents schema-based structured output for Gemini, while Microsoft documents structured outputs in its Agent Framework. Snowflake’s Cortex Code Agent SDK also documents structured output patterns. These mechanisms differ in supported schema subsets and failure behavior, so check the provider’s current documentation before relying on a keyword such as “strict.”
Minimal Python extraction shape
from dataclasses import dataclass
from typing import Optional
@dataclass
class TaskRecord:
summary: str
assignee: Optional[str]
priority: Optional[str]
due_date: Optional[str]
labels: list[str]
needs_clarification: bool
evidence: list[str]
# The model call is provider-specific. The important contract is that
# the response is parsed against the schema before this object is used.
def accept_record(parsed: TaskRecord) -> TaskRecord:
allowed = {None, "low", "medium", "high", "urgent"}
if parsed.priority not in allowed:
raise ValueError("priority is outside the allowed enum")
if parsed.due_date is not None and len(parsed.due_date) != 10:
raise ValueError("due_date must be YYYY-MM-DD")
return parsed
SDK parsing can turn a successful response into a native object, but your application still needs checks such as date validity, identifier existence, authorization, and grounding. A parser should reject malformed data; it cannot determine whether “Friday” was interpreted correctly for your business.
4. Validate structure, semantics, and grounding
Run validation in layers:
- Transport validation: confirm the model call completed and the response is not a refusal, truncation, or tool error.
- Schema validation: parse JSON and reject unknown keys, wrong types, missing required properties, and invalid enum values.
- Semantic validation: check ISO dates, identifier formats, allowed label sets, and domain rules such as “a completed task cannot have a future completion date.”
- Grounding validation: verify that populated values are supported by the source description. Evidence strings help reviewers, but do not replace source comparison.
- Workflow validation: confirm that the current user is allowed to create, assign, or update the resulting record.
Keep the original description, schema version, model configuration, parsed record, validation errors, and final disposition together. This makes corrections reproducible when a prompt or schema changes.
Dates and relative language
Relative dates are a frequent source of silent errors. Supply the reference date and timezone as explicit context, and require the agent to return an ISO date only when the wording resolves unambiguously. If your system cannot establish the intended date, return null and ask a follow-up question. Do not let a downstream scheduler interpret a vague date differently from the extractor.
5. Build explicit failure paths
Production code should treat these outcomes differently:
| Outcome | Recommended action |
|---|---|
| Schema validation failure | Retry with the same source and a repair instruction, then quarantine after a limit |
| Missing required information | Return a clarification request instead of inventing a value |
| Ambiguous wording | Set a review flag and preserve the ambiguity in evidence |
| Unsupported inference | Drop the value, record a validation error, and inspect the instruction |
| Refusal, timeout, or rate limit | Apply bounded retry policy and keep the task uncommitted |
| Conflicting source statements | Escalate or apply a documented precedence rule |
A repair retry should not ask the model to “try harder” without context. Include the validation error, the original task, and the same schema. Limit retries so malformed output cannot create an infinite loop or duplicate side effect.
6. Evaluate extraction quality with representative tasks
Build a small labeled set from real descriptions before selecting a platform or prompt. Include short and long tasks, missing fields, colloquial language, conflicting instructions, multiple dates, identifiers, and adversarial text that attempts to change the extraction rules.
Track separate measures:
- Schema success: percentage of responses that parse and validate.
- Field precision: how often populated values are supported and correct.
- Field recall: how often relevant stated values are captured.
- Unsupported inference rate: values present in output without source support.
- Clarification quality: whether genuinely ambiguous cases are flagged.
- Operational metrics: latency, token usage, retries, and failure cost.
The documentation reviewed for OpenAI, Google, Microsoft, and Snowflake describes structured-output mechanisms, not a controlled accuracy benchmark for this exact task. Compare candidates on the same examples, schema, error definitions, and operating conditions. Schema conformance alone is not evidence that one provider extracts facts more accurately.
7. Add tools without losing the final schema
An agent may need tools to resolve an assignee ID, look up a project, or fetch a timezone. Keep tool arguments and the final record separate. Tool results should be treated as external data that also needs validation. The final response should still conform to the declared record schema, with an evidence trail showing whether a value came from the task text or an authorized lookup.
Use a two-phase design for consequential actions: first extract and validate a proposed record; then ask for confirmation or apply deterministic authorization checks before creating tickets, sending messages, or changing production data. Never let a schema-valid object bypass permissions.
8. Troubleshooting common errors
“The response is JSON, but parsing still fails”
Cause: markdown fences, trailing commentary, truncated output, or a provider-specific refusal object. Fix: use the provider’s structured-output mode, inspect the raw response type, enforce a maximum output size, and reject anything outside the expected envelope.
“Required fields are always filled with guesses”
Cause: the prompt implies every field must have a value, or the schema does not allow null. Fix: allow nullable fields, define absent versus ambiguous behavior, and include examples where the correct answer is null.
“Dates are off by one day”
Cause: timezone conversion or an unstated reference date. Fix: pass an explicit timezone and current date, require ISO output, and validate the date before scheduling.
“Enum values vary in capitalization”
Cause: the model is not constrained to the enum or application code normalizes inconsistently. Fix: enforce the enum in the schema and reject values outside it; do not silently map unknown values to a nearby category.
“The agent misses a detail buried in a long task”
Cause: unclear field definitions, competing instructions, or context limits. Fix: split very long descriptions into bounded sections, ask for evidence per field, and add representative long examples to evaluation. If splitting, reconcile records deterministically and flag conflicts.
“A valid record triggers the wrong action”
Cause: structural validation was treated as proof of correctness or authorization. Fix: add semantic, grounding, and permission checks, then require confirmation for irreversible operations.
9. Performance, reliability, and cost considerations
Schema complexity affects latency and output size. Keep descriptions concise, avoid redundant fields, and use enums instead of open-ended prose where the domain permits. Evidence improves reviewability but increases tokens; require short evidence strings rather than full restatements.
Use bounded retries with exponential backoff for transient provider errors, but do not retry deterministic schema or permission failures unchanged. Add idempotency keys around downstream writes so a retry cannot create duplicate tickets. Log request IDs, schema versions, and validation outcomes without storing sensitive task text unless your retention policy allows it.
Estimate cost from input tokens, output tokens, tool calls, and retry frequency. A cheaper model may be suitable for simple classification, while complex multi-field extraction needs evaluation before you trade quality for price. Measure p50 and tail latency separately, because a workflow that is fast on average can still block users during retries or tool calls.
10. Or skip the browser setup
If your agent needs a screenshot of a task-related web page as additional context, ScreenshotNeo provides a single-call website screenshot API and MCP server. The DIY approach is to configure a browser, load the page, wait for content, handle consent banners, and save an image. ScreenshotNeo handles that capture flow through its API.

cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. You get 1,000 screenshots a month free with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
11. Short FAQ
Can an AI agent convert natural language directly to JSON?
Yes, when the application supplies a schema and parses the result. Plain prompting alone does not guarantee valid or complete data.
Should missing fields be empty strings?
Usually no. Use null for an unknown scalar and an empty array for a known collection with no entries. Document the convention in the schema.
Is strict structured output enough for production?
No. It improves shape guarantees, but you still need semantic, grounding, authorization, and workflow checks.
How many examples should the prompt include?
Use enough to demonstrate ambiguous and missing cases. A small, representative set is more useful than many near-duplicates; measure the effect on your evaluation set.
When should a human review the result?
Route records with contradictions, unresolved ambiguity, unsupported values, failed validation, or consequential side effects to a review step.