ScreenshotNeo

BlogHow-to

UTM Validator: Check Campaign Tracking URLs

Validate UTM links before launch: check required fields, duplicates, naming, encoding, redirects, and why a structurally valid URL can still miss GA4 attribution.

By the ScreenshotNeo team29 September 202612 min read

UTM Validator: Check Campaign Tracking URLs

A UTM validator checks that a campaign tracking URL is structurally sound and follows your naming rules before you publish it. At minimum, verify that the destination has a valid scheme and host, that utm_source, utm_medium, and utm_campaign are present and non-empty, and that there are no duplicate tracking parameters. Then test the final redirect and landing page: a URL can pass validation while a redirect drops its query string or analytics collection never runs.

Google recommends using those three core parameters whenever you add campaign parameters. UTM values are case-sensitive, so utm_source=Meta and utm_source=meta can be recorded as different values. A validator is therefore both a URL check and a campaign-governance check. Google Analytics Help: URL builders and campaign parameters

1. What a UTM validator should check

A useful validator does more than look for the substring utm_. It parses the URL into its destination and query parameters, applies required-field rules, checks values, and reports issues in a way a marketer or developer can fix.

A validator parses the destination and each parameter separately so it can catch missing fields and duplicates.
A validator parses the destination and each parameter separately so it can catch missing fields and duplicates.
Check What to validate Why it matters
Destination Absolute URL with an allowed scheme such as https and a hostname Prevents publishing a malformed or unintended link
Core fields utm_source, utm_medium, utm_campaign exist and are not blank These are the fields Google recommends for tagged URLs
Duplicates Each UTM key appears once Duplicate values can produce ambiguous or unexpected attribution
Naming policy Approved spelling, case, separators, source and medium vocabulary Consistent values avoid splitting reports across near-duplicates
Optional detail utm_content, utm_term, and other fields are used consistently Supports creative and paid-keyword differentiation where needed
Runtime behavior Redirects preserve the query and the landing page loads analytics Syntax alone cannot prove collection or attribution will work

Google documents these campaign fields: utm_id, utm_source, utm_medium, utm_campaign, utm_source_platform, utm_term, utm_content, utm_creative_format, and utm_marketing_tactic. The last two are not currently reported in some Analytics properties, so confirm your reporting setup before making them required. Field definitions and guidance

2. A practical UTM validation checklist

  1. Parse the URL. Confirm there is a valid scheme and host. Decide whether your policy permits http; for public campaign links, many teams require https.
  2. Require the core fields. Check source, medium, and campaign. Treat missing and whitespace-only values as empty.
  3. Require an ID when the workflow needs it. Campaign-data imports for non-Google campaign URLs require utm_id along with source, medium, and campaign. This is not a universal requirement for every tagged link. Google Analytics Help: import campaign data
  4. Detect duplicate keys. Do not silently pick the first or last value. Flag repeated UTM parameters for correction, including a second set appended by a link builder.
  5. Apply a naming dictionary. Validate allowed source and medium values, capitalization, campaign pattern, spelling, and separators. Compare against exact approved values rather than fuzzy matching.
  6. Review optional fields. Use utm_content for distinctions such as creative variants and utm_term for paid keyword identification when those distinctions are useful to your reporting.
  7. Resolve dynamic macros appropriately. Preserve platform placeholders only when the ad platform will substitute them and the destination can keep the resulting query. For campaign-data import, Google advises resolving dynamic values at the advertising-platform level. Import requirements
  8. Test the final destination. Follow redirects and inspect the landing URL. Confirm the query survives and that the page’s analytics implementation and consent behavior are appropriate.

3. Build a small runnable validator in Python

This standard-library script validates a single URL, reports duplicate UTM keys, checks required values, optionally enforces exact approved sources and media, and prints a machine-readable result. It does not make a network request, follow redirects, or verify analytics collection.

#!/usr/bin/env python3
import json
import sys
from urllib.parse import parse_qsl, urlsplit

REQUIRED = ("utm_source", "utm_medium", "utm_campaign")
# Replace these examples with your organization's canonical values.
ALLOWED = {
    "utm_source": {"newsletter", "google", "meta"},
    "utm_medium": {"email", "cpc", "paid_social"},
}


def validate(raw_url):
    errors = []
    warnings = []
    try:
        parts = urlsplit(raw_url)
        # Accessing .port can raise ValueError for an invalid port.
        _ = parts.port
    except ValueError as exc:
        return {"valid": False, "errors": [f"Malformed URL: {exc}"], "warnings": []}

    if parts.scheme.lower() not in {"http", "https"}:
        errors.append("Destination must use http or https")
    if not parts.hostname:
        errors.append("Destination must include a hostname")

    pairs = parse_qsl(parts.query, keep_blank_values=True)
    values = {}
    for key, value in pairs:
        if key.lower().startswith("utm_"):
            values.setdefault(key, []).append(value)

    for key, entries in values.items():
        if len(entries) > 1:
            errors.append(f"Duplicate parameter: {key}")

    # Parameter names are treated case-insensitively here to catch spelling
    # variants, while values retain their case because Google values are case-sensitive.
    normalized = {}
    for key, entries in values.items():
        normalized.setdefault(key.lower(), []).extend(entries)

    for key in REQUIRED:
        entries = normalized.get(key, [])
        if not entries or not entries[0].strip():
            errors.append(f"Missing or empty required parameter: {key}")

    for key, approved in ALLOWED.items():
        entries = normalized.get(key, [])
        if entries and entries[0] not in approved:
            errors.append(f"Unapproved value for {key}: {entries[0]!r}")

    if "utm_id" not in normalized:
        warnings.append("Add utm_id if required by campaign import or internal policy")
    if parts.fragment:
        warnings.append("Check that the fragment is intentional; it is not sent to the server")

    return {
        "valid": not errors,
        "destination": f"{parts.scheme}://{parts.netloc}{parts.path or '/'}",
        "parameters": values,
        "errors": errors,
        "warnings": warnings,
    }


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit(f"Usage: {sys.argv[0]} 'https://example.com/?utm_source=...'")
    print(json.dumps(validate(sys.argv[1]), indent=2))

Save it as validate_utm.py and run it with a quoted URL:

python3 validate_utm.py 'https://example.com/pricing?utm_source=newsletter&utm_medium=email&utm_campaign=fall_launch'

Use parse_qsl rather than converting query pairs directly to a dictionary: a dictionary would discard repeated keys and make duplicate detection impossible. The sample treats parameter names case-insensitively for duplicate and required-field checks, but checks values exactly. Adapt the policy to your conventions; do not normalize values automatically if exact capitalization is important to your reports.

4. Parse and validate with JavaScript

In a browser or Node.js, the built-in URL and URLSearchParams APIs handle decoding and preserve access to repeated parameters through getAll. This example runs in Node.js without dependencies.

// validate-utm.mjs
const REQUIRED = ['utm_source', 'utm_medium', 'utm_campaign'];
const allowed = {
  utm_source: new Set(['newsletter', 'google', 'meta']),
  utm_medium: new Set(['email', 'cpc', 'paid_social']),
};

function validate(raw) {
  const errors = [];
  const warnings = [];
  let url;
  try {
    url = new URL(raw);
  } catch {
    return { valid: false, errors: ['Malformed absolute URL'], warnings };
  }
  if (!['http:', 'https:'].includes(url.protocol)) {
    errors.push('Destination must use http or https');
  }
  const keys = [...new Set([...url.searchParams.keys()])];
  const normalized = new Map();
  for (const key of keys) {
    if (!key.toLowerCase().startsWith('utm_')) continue;
    const values = url.searchParams.getAll(key);
    if (values.length > 1) errors.push(`Duplicate parameter: ${key}`);
    const lower = key.toLowerCase();
    normalized.set(lower, [...(normalized.get(lower) ?? []), ...values]);
  }
  for (const key of REQUIRED) {
    const values = normalized.get(key) ?? [];
    if (!values.length || !values[0].trim()) {
      errors.push(`Missing or empty required parameter: ${key}`);
    }
  }
  for (const [key, approved] of Object.entries(allowed)) {
    const value = normalized.get(key)?.[0];
    if (value !== undefined && !approved.has(value)) {
      errors.push(`Unapproved value for ${key}: ${value}`);
    }
  }
  if (!normalized.has('utm_id')) {
    warnings.push('Add utm_id if required by campaign import or internal policy');
  }
  return {
    valid: errors.length === 0,
    destination: `${url.origin}${url.pathname}`,
    parameters: Object.fromEntries(normalized),
    errors,
    warnings,
  };
}

const result = validate(process.argv[2] ?? '');
console.log(JSON.stringify(result, null, 2));
process.exitCode = result.valid ? 0 : 1;
node validate-utm.mjs 'https://example.com/?utm_source=meta&utm_medium=paid_social&utm_campaign=fall_launch'

5. Quick checks with cURL and Python requests

For a one-off syntax check, pass the URL to a local parser. cURL can also show a redirect chain, but it does not understand your naming dictionary or determine whether analytics recorded a visit.

curl -sS -L -D response-headers.txt -o /dev/null \
  -w 'final_url=%{url_effective}\nstatus=%{http_code}\nredirects=%{num_redirects}\n' \
  'https://example.com/?utm_source=newsletter&utm_medium=email&utm_campaign=fall_launch'

Review response-headers.txt and the final URL for a redirect that strips query parameters. Avoid sending private or sensitive campaign tokens to third-party inspection services. For a Python network check using the commonly installed requests package:

import requests

url = "https://example.com/?utm_source=newsletter&utm_medium=email&utm_campaign=fall_launch"
response = requests.get(url, allow_redirects=True, timeout=20)
print("status:", response.status_code)
print("final URL:", response.url)
print("redirect count:", len(response.history))
for hop in response.history:
    print(hop.status_code, "->", hop.headers.get("location"))
response.raise_for_status()

A successful HTTP response confirms only that a request reached a server. It does not prove the browser executed JavaScript analytics, consent allowed collection, or campaign attribution appeared in a report.

6. Naming rules that prevent reporting splits

Write down a small canonical dictionary before adding more validation logic. For each team, define the allowed source values, medium values, campaign pattern, casing rule, and optional content and term conventions. For example, a lowercase-only policy might permit utm_source=meta and reject Meta; the important part is consistency and exact matching.

  • Use one spelling for each platform and channel. Do not alternate between abbreviations and full names without a rule.
  • Choose separators deliberately, such as underscores between campaign components, and document the expected pattern.
  • Keep values stable across link builders and manual workflows. Do not “fix” capitalization on the fly without checking existing reporting conventions.
  • Use campaign IDs where teams need a stable identifier even when a human-readable campaign name changes.
  • Validate the decoded value, but preserve correct URL encoding when sharing the link. Spaces and reserved characters should be percent-encoded by a URL builder.

Google notes that values are case-sensitive and advises consistent values for platforms, channels, campaign names, content, and terms. For campaign-data imports, imported values must match the values logged by Analytics, including capitalization. Campaign import matching rules

7. Why a valid UTM URL can still fail in GA4

Validation confirms the URL’s structure and your conventions. It cannot guarantee attribution. Diagnose the remaining path separately:

A URL can pass syntax checks while a redirect still strips its campaign parameters.
A URL can pass syntax checks while a redirect still strips its campaign parameters.
  • A redirect removed the query. Inspect each redirect hop and configure the redirecting site to preserve the query string.
  • The landing page does not collect analytics. Confirm the correct Analytics property and tag load on the final page.
  • Consent or browser controls prevent collection. Test with the site’s intended consent state and browser environment.
  • Auto-tagging affects traffic-source data. Google explains that when GCLID or DCLID cannot be used as intended, Analytics derives cross-channel traffic-source dimensions from available UTM parameters. Check the actual collected dimensions and tagging configuration. Google Analytics traffic-source dimensions
  • Campaign values differ in case or spelling. Compare the exact URL values with the values used in reports and imported campaign data.
  • A macro remains unresolved. Check whether the advertising platform substitutes it and whether the final click URL retains the resulting value.

8. Workflow options: one-off checks or governance

For a small team, a local script and a shared naming sheet can be enough. Larger publishing workflows benefit from a central validator that applies a versioned dictionary and records who approved a link. Compare tools on required-field rules, duplicate detection, custom naming dictionaries, dynamic-macro handling, redirect and landing-page inspection, bulk or API workflows, governance logs, export options, and privacy treatment.

Keep validation close to link creation: in a campaign builder, content-management form, spreadsheet import, or pre-launch checklist. Return actionable messages such as “utm_medium must be one of email, cpc, paid_social” instead of a generic failure. Store the canonical rules in one maintained location, and make batch validation report the row or campaign that needs correction.

9. Troubleshooting common validation errors

Symptom Likely cause Fix
“Missing required parameter” Parameter absent, misspelled, or present with an empty value Use the exact key and provide a non-empty value
“Duplicate parameter” A builder appended UTMs to an already-tagged destination Keep one authoritative value per key; remove the old set or merge intentionally
“Unapproved value” Different capitalization, alias, or spelling from the naming dictionary Use the canonical value; review reporting implications before changing established names
Malformed URL Missing scheme/host, invalid port, or unescaped characters Use an absolute URL and encode query values with a URL library
UTMs disappear after a click Redirect or routing rule drops the query string Inspect every hop and configure query preservation
URL passes but GA4 has no campaign Tag absent, consent blocks collection, wrong property, or reporting delay/configuration issue Verify the final page and analytics collection independently; inspect the exact collected traffic-source values
Campaign import rejects values Required ID missing or imported values do not exactly match logged values Include utm_id where required and match capitalization and spelling exactly

10. Use a screenshot to inspect the final landing page

When a link reaches a page but the rendered state is unclear, capture the final page after following the redirect. A screenshot can help spot a broken landing page, consent overlay, or missing visible content. It is a visual check, not proof that GA4 received an event; inspect analytics separately.

Or skip the browser setup

For a rendered landing-page check, ScreenshotNeo is a website screenshot API and MCP server. One GET request returns an image or PDF. Its consent handling accepts the banner like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome indicated by X-Page-Verdict and X-Billed. AI agents can use its MCP tools, including take_screenshot, get_page_info, and capture_pdf. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com/pricing \
  -o landing.webp

See the ScreenshotNeo API documentation for request options. Create a free account for 1,000 screenshots a month with no card: sign up for ScreenshotNeo.

11. Performance, reliability, and cost

Local parsing is fast, deterministic, and incurs no request cost. It is the right first gate for large batches because it catches syntax, duplicates, and naming violations without loading destination pages. Network checks add latency and can fail for reasons unrelated to UTM correctness, including site availability, rate limits, bot defenses, and slow scripts. Set explicit timeouts, limit concurrency, and avoid retrying non-idempotent workflows unnecessarily.

Run redirect and rendered-page checks on a sample or on links where risk justifies the extra work. Record the final URL, response status, and validation-rule version, but avoid logging secrets embedded in URLs. Treat campaign URLs as public once distributed; do not put personal data or confidential identifiers in UTM values. Keep a clear distinction in tooling between “URL passes policy” and “landing page reached”; neither state by itself guarantees Analytics attribution.

12. Frequently asked questions

No. Google recommends source, medium, and campaign for tagged URLs. Campaign-data imports for non-Google URLs have additional requirements that include utm_id; use it when that workflow or your own governance requires it.

Should a validator automatically lowercase values?

Usually it should report a violation and let the owner correct the source link. Automatic rewriting can change a value that must exactly match existing reports or an import.

Can a validator guarantee a campaign appears in GA4?

No. It can establish structure and naming compliance. Redirect behavior, page instrumentation, consent, browser conditions, and traffic-source configuration also affect collection and reporting.

Should I keep UTM parameters when redirecting?

Preserve them when the destination needs them for attribution, and verify the final URL rather than assuming each redirect does so.

Sources