BREAKBOT / FIELD MANUAL

From endpoint
to evidence.

A short guide to the implemented CLI. BreakBot runs locally; this website presents the project and does not provide a hosted scanner or public server database.

01 / Set up the CLI

Requirements: Python 3.12+, uv, and PostgreSQL 16 for persistence and index commands. Probing a single endpoint does not require PostgreSQL or AI dependencies.

The repository is currently private. GitHub access must be granted before cloning it.

git clone https://github.com/Vadyanik/breakbot.git
cd breakbot
uv sync --locked --dev
uv run breakbot --help

Use a server you operate or have permission to check:

uv run breakbot probe <host:port>
uv run breakbot probe <host:port> --json
uv run breakbot probe <host:port> --layout classic

The default output is a Rich card. --layout classic provides section-by-section text. Optional favicon rendering uses a compatible Ghostty or Kitty terminal.

02 / Build and query your index

The repository includes a Docker Compose configuration for a localhost PostgreSQL database. Copy the example environment and replace its placeholder credentials before starting it.

cp .env.example .env
# Edit .env: replace change-me in both password and DATABASE_URL.
docker compose up -d --wait
set -a
source .env
set +a
uv run alembic upgrade head
uv run breakbot probe <host:port> --save

Keep the database volume when it contains observations. If PostgreSQL already exists, configure DATABASE_URL and apply migrations to that database instead.

uv run breakbot scan targets.txt --concurrency 32
uv run breakbot import masscan.json --concurrency 32
uv run breakbot search --version 1.21 --players-min 5
uv run breakbot search --software Paper --mod-loader Forge --json
uv run breakbot server <host:port> --raw
uv run breakbot rescan --older-than 3600
uv run breakbot stats

Target files contain one host or host:port per line. import consumes external masscan JSON and verifies open TCP candidates with Minecraft Status. It does not run masscan itself. Search filters combine version, explicitly reported software, loader, player range, MOTD, and sourced country metadata. The read commands support JSON output.

Rechecks preserve the last successful state after failure. The age of a stored result matters: server distinguishes the last check from the last successful response.

03 / Optional AI-assisted Server Classification

In development. The explicit analysis command, Claude client, validators, caching, and separate storage are implemented locally. This website audit verified 63 focused tests; 3 database tests were skipped. No live Claude response was verified. AI-tag filters, AI display in server, and automatic enrichment remain planned.

Analysis starts from a previously saved successful Status response. It does not probe the server. New analysis requires an Anthropic API key and can incur provider charges.

uv sync --extra ai --dev
uv run alembic upgrade head
# Set ANTHROPIC_API_KEY securely in your process environment.
# DATABASE_URL must point to your migrated PostgreSQL database.
uv run --extra ai breakbot analyze <host:port>
uv run --extra ai breakbot analyze <host:port> --json

Supported options are --json, --force, and --model. --force bypasses the cache and requests a fresh Claude analysis and may incur API charges. A matching validated cache can be read without an API key. Model, input hash, prompt version, and schema version determine cache eligibility.

Structured outputWhat it means
game_modesControlled vocabulary; each mode carries advertised or inferred basis and supporting excerpts.
languageAdvertised or inferred language; unknown is null.
regionAdvertised region only. Does not establish physical location.
tagsGameplay tags with basis and evidence. Stored classifications, not currently searchable CLI filters.
summaryAI interpretation, limited to 400 characters, or null.
claimsTyped advertised claims, such as economy, crossplay, or land claims. They remain unverified.
uncertainExplicit ambiguity and conflicting claims.

Results live in ai_enrichments, separate from scanner state, with sanitized input, requested and actual model, input hash, prompt/schema versions, provider usage when available, and observation/analysis timestamps. Analysis may use old metadata after failed rescans.

Ordinary scans, rescans, search, and server viewing do not request Claude. Model output cannot overwrite reported scanner fields. Evidence checks require the quotation to occur in the submitted input; they do not verify the model’s interpretation or the server’s claim.

Input and privacy limits

The input builder allowlists cleaned MOTD, version, explicitly reported software/proxy/loader, and mod metadata. MOTD is capped at 4,096 characters, text metadata at 256 characters, mod entries at 64, and the serialized input at 24 KiB. Player samples and counts, endpoints, raw Status payloads, and arbitrary fields are omitted from the structured input.

Embedded identifiers are redacted on a best-effort basis. Arbitrary server prose can still contain personal information. Redaction is not a guarantee of anonymity. Server text is untrusted input, and schema/evidence validation reduces risk without guaranteeing prompt-injection resistance. AI interpretations may be inaccurate.

04 / Where the demonstrations come from

The homepage uses the repository’s tests/fixtures/normal.json. A localhost TCP responder supplied that fixture to the actual breakbot probe command. The displayed endpoint and latency are capture-specific; they do not describe a public server or benchmark.

{
  "version": {
    "name": "1.21.8",
    "protocol": 772
  },
  "players": {
    "online": 7,
    "max": 20
  },
  "description": {
    "text": "Private Survival"
  },
  "enforcesSecureChat": true,
  "previewsChat": false
}

The classification example is hand-composed illustrative data validated by BreakBot’s actual schema and evidence validator, then passed to its real CLI renderer. No Claude request was made. It is intentionally narrow: “Survival” supports an advertised game mode; language, region, and authentication properties remain unknown.

The website does not perform scans or AI analysis. Copy controls copy text; inspection controls reveal static evidence and architecture notes. There is no simulated live activity.

05 / Current development state

The core V1 CLI stages are implemented: probe, parse, save, scan, import, search, and rescan. V2 plans cover continuous discovery and scheduling, reverse player lookup, separate Login checks, throughput work, operational retention, and owner opt-out. The roadmap has no claimed delivery dates.

This content was audited against the current local checkout on October 8, 2026, including uncommitted Claude integration work. It does not claim that all local changes have been released on GitHub. A release is identified by a version tag.

For the latest source and configuration details, see the private project repository and its README, development plan, and docs/ai-classification.md. When publishing the site, publish or reconcile its supporting documentation too.

Read collection and responsible-use notes →