ScreenshotNeo

BlogHow-to

What Is Static Code Analysis and How to Set It Up?

Learn what static code analysis checks, how SAST works, and how to set up practical local, pull-request, and CI quality gates.

By the ScreenshotNeo team1 October 20267 min read

Static code analysis examines source code, bytecode, or binaries without running the application. It can catch defects and security problems before deployment, in an IDE, pre-commit hook, pull request, or CI job. A reliable setup combines fast pattern rules with deeper data-flow analysis, a buildable project, triage, and a quality gate.

This guide shows how to choose an analyzer, run a baseline scan, integrate GitHub Actions, publish SARIF results, tune rules, and understand the limits of SAST.

1. What static code analysis checks

Static analysis tools inspect code using lexical and pattern matching, abstract-syntax-tree (AST) matching, data-flow analysis, and taint analysis. Taint analysis follows user-controlled data from a source, such as an HTTP parameter, to a dangerous sink, such as a SQL query or shell command, and reports flows that are not safely sanitized.

OWASP describes static analysis as usually being performed as part of code review during the implementation phase of a secure development lifecycle (OWASP Static Code Analysis).

Analysis depth Typical findings Trade-off
Lexical and pattern rules Bad APIs, insecure functions, formatting, obvious secrets Fast and easy to run; limited context
AST rules Framework misuse, dangerous syntax, missing checks More precise structure-aware matching
Data-flow and taint analysis Unsanitized input reaching SQL, HTML, files, or commands More compute, configuration, and build context

2. Static analysis, code review, and dynamic testing

SAST is preventive feedback. It does not execute the application and cannot reliably discover every authorization, authentication, cryptography, configuration, dependency, or external-system problem. False positives and false negatives are normal.

  • Peer review: evaluates intent, architecture, business rules, and maintainability.
  • Dependency and secret scanning: checks third-party packages, licenses, and exposed credentials.
  • Dynamic testing (DAST and integration tests): observes a running system and its runtime behavior.
  • Infrastructure and configuration checks: inspect deployment manifests, cloud permissions, and server settings.
  • Runtime monitoring: detects issues after release.

Use static analysis as one control in that wider assurance process, not as a replacement for it.

3. Choose a tool and analysis depth

Choose based on language and framework coverage, pattern versus data-flow depth, build requirements, IDE and CI integration, SARIF support, licensing, scan time, customization, and false-positive handling.

Tool Useful starting point Best fit
GitHub CodeQL Semantic analysis integrated with GitHub code scanning; default or advanced setup Repositories hosted on GitHub that need deep analysis and pull-request alerts
Semgrep Fast pattern rules, registry, broad language support, and CI/AppSec workflows Teams that want quick custom rules and incremental adoption
SonarQube Hosted or self-managed automated code review with pull-request decoration Organizations wanting a centralized quality platform
Language-specific analyzers Examples include Bandit (Python), Brakeman (Ruby on Rails), gosec (Go) Framework-specific security checks with minimal setup

A practical sequence is: start with the vendor’s defaults, add a language-specific analyzer, then add interprocedural or taint rules for sensitive boundaries.

4. Inventory the repository before scanning

  1. List production languages, generated code, build systems, package managers, and test commands.
  2. Identify HTTP input, database queries, file operations, deserialization, template rendering, shell execution, authentication, and authorization boundaries.
  3. Record generated, vendored, test, and third-party directories so they can be included or excluded intentionally.
  4. Make the project buildable with a clean checkout. Missing libraries and incomplete build steps reduce coverage for many analyzers.

5. Run a first local baseline

Install one analyzer for your main language, run its default rules, export findings, and classify each result as real, false positive, duplicate, or accepted risk. Do not fail the whole team on pre-existing alerts before triage.

Semgrep quick start

python -m pip install semgrep
semgrep scan --config=auto --json --output=semgrep-baseline.json .

Python Bandit quick start

python -m pip install bandit
bandit -r . -f json -o bandit-baseline.json

CodeQL command-line outline

# Install the CodeQL CLI from GitHub's releases, then:
codeql database create codeql-db --language=javascript --command="npm ci && npm run build"
codeql database analyze codeql-db javascript-security-and-quality.qls \
  --format=sarif-latest --output=codeql-results.sarif

Use the language pack and query suite that match your repository. A compiled project usually needs a complete build command; interpreted languages may need their dependency installation and test setup.

6. Add pull-request and CI scanning

GitHub CodeQL default setup

GitHub CodeQL can scan pushes, pull requests, and weekly scheduled builds for supported languages. Enable it under your repository’s Security settings, select default or advanced setup, and verify that the detected languages and build mode are correct.

Semgrep in GitHub Actions

name: static-analysis
on:
  pull_request:
  push:
    branches: [main]

jobs:
  semgrep:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      security-events: write
    steps:
      - uses: actions/checkout@v4
      - uses: semgrep/semgrep-action@v1
        with:
          config: p/default
          generateSarif: true

Pin action versions according to your organization’s policy and review the action’s current configuration options in its documentation.

Upload SARIF results

GitHub code scanning accepts SARIF 2.1.0 results from CodeQL and third-party analyzers. A generic upload step is:

- name: Upload SARIF
  uses: github/codeql-action/upload-sarif@v3
  with:
    sarif_file: semgrep-results.sarif

Ensure the job has security-events: write permission and that the SARIF file exists at the path supplied. See GitHub’s SARIF upload documentation.

7. Define a quality gate that teams can keep

  1. Block new high-severity findings on changed code.
  2. Require a justification, owner, and expiry or review date for suppressions.
  3. Keep lower-severity and legacy findings visible as a backlog.
  4. Allow emergency overrides only with an auditable issue or pull-request comment.
  5. Measure rule noise, scan duration, and remediation age rather than chasing a raw alert count.

A changed-code gate lets a team improve safely without making an old backlog an excuse to ignore new defects.

8. Configure rules and exclusions

  • Include production code: exclude generated and vendored files unless they are part of your threat model.
  • Model frameworks: teach the analyzer which functions read requests, validate permissions, build queries, or sanitize output.
  • Set severity deliberately: reserve blocking status for issues with a clear remediation path.
  • Use narrow suppressions: suppress a specific rule at the smallest scope and record why it is safe.
  • Review configuration as code: version rule files and require review for changes.

Overly broad exclusions hide defects; overly broad rule sets create alert fatigue. Tune from real findings and revisit coverage when languages, frameworks, or dependencies change.

9. Troubleshooting common failures

Symptom Likely cause Fix
No findings Wrong language, empty path, or unsupported framework Check detected languages, scan paths, and analyzer support; run a known vulnerable test fixture.
Build or database creation fails Missing dependencies, generated files, or incorrect build command Run the same clean build locally, install dependencies, and provide the exact build command.
Thousands of alerts Legacy code or an overly broad ruleset Save a baseline, gate only new high-severity findings, then triage by ownership and risk.
False positives Analyzer lacks framework or project context Add models, narrow the rule, or document a scoped suppression.
Pull-request annotations missing SARIF path, permissions, or upload step is wrong Confirm the file exists, use SARIF 2.1.0, grant security-events: write, and inspect the workflow log.
CI is too slow Full history, large generated trees, or deep taint analysis on every change Cache dependencies, exclude generated output, scan changed code on PRs, and schedule full scans.
Results differ locally and in CI Different analyzer versions, rules, build flags, or dependency locks Pin versions, commit configuration, and use the same lockfile and build command.

10. Performance, reliability, and cost planning

  • Fast feedback: run pattern and AST rules on every commit or pull request.
  • Deep coverage: run interprocedural and taint analysis with a complete build, often on pull requests and scheduled jobs.
  • Reliable results: pin tool and rule versions, cache dependencies, preserve SARIF artifacts, and alert when a scan itself fails.
  • Resource control: split jobs by language, exclude generated output, and use incremental or changed-file modes where supported.
  • Cost: account for CI minutes, hosted analyzer plans, storage, and developer triage time. A noisy gate can cost more in review time than the scanner costs to run.

There is no universal scan-time or detection-rate number. Measure your own repository after establishing a baseline, and treat a passed scan as evidence that configured rules ran, not proof that the code is secure.

11. Or skip the browser setup

If your engineering workflow also needs screenshots of build reports, documentation, or rendered pages, ScreenshotNeo provides a website screenshot API and MCP server. It removes cookie banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; and an MCP server lets AI agents take screenshots.

One request returns PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for the 63 capture options, response headers, and async workflows. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

12. FAQ

Can static analysis prove that code is secure?

No. It finds patterns and flows covered by its rules. Combine it with review, dependency and secret scanning, dynamic tests, configuration checks, and runtime monitoring.

Should every finding fail the build?

No. Start by blocking new high-severity findings and require documented suppressions. Keep lower-severity and legacy findings visible.

Do analyzers need the application to run?

No. They inspect code or compiled artifacts without executing the application, although many need dependencies and a complete build for accurate results.

When should a full scan run?

Run fast checks on pull requests and schedule deeper full-repository scans, especially after dependency, framework, or analyzer changes.

What is SARIF?

SARIF is a standardized results format that lets tools such as CodeQL and Semgrep publish findings to GitHub code scanning and other compatible systems.