| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
| Name | Name | Last commit date | ||
|---|---|---|---|---|
PIIFlow is a policy-aware static taint analyzer for Python that traces sensitive data into LLM API calls and explains whether each flow violates a declared privacy policy.
A sensitive flow is not necessarily a policy violation. PIIFlow first detects the flow, then evaluates it against provider, trust-zone, category, and sanitizer requirements.
Python LLM applications often move request data through helpers, prompt builders, containers, and SDK calls. A code review may miss that an email address, diagnosis, account number, or government identifier eventually reaches an external model provider. PIIFlow scans source code without executing it, reports the source-to-sink path, and can apply a YAML privacy policy to decide whether a detected flow is compliant or a violation.
PIIFlow is a bounded developer and research prototype. It is not runtime data loss prevention, a legal compliance engine, or a guarantee that every Python data flow is found.
def build_prompt(email: str) -> str:
return f"Summarize the customer account for {email}"
def handle_email(request, openai_client) -> None:
email = request.json["email"]
prompt = build_prompt(email)
openai_client.responses.create(input=prompt)Run:
uv run piiflow scan examples/vulnerable_appShortened output from the actual CLI:
PIIFlow 0.3.0
PIIFLOW001 HIGH Email data may reach an external LLM.
Source: app.py:10 request.json['email']
Path: request.json['email'] -> email -> build_prompt(email)
Sink: app.py:11 openai_client.responses.create(input=)
Remediation: Apply an approved sanitizer before the sink, remove unnecessary
sensitive fields, use an internal/local model if policy allows, or update rules
only when this flow is intentionally permitted.
Policy:
version: 1
default_action: deny
rules:
- id: require-redacted-email-openai
description: Email may be sent to OpenAI only after redact_email.
categories: [EMAIL]
providers: [openai]
trust_zones: [external]
action: require_sanitizer
sanitizers: [redact_email]Run:
uv run piiflow scan examples/policy_apps/redacted_email.py \
--policy examples/policies/sanitized_email.ymlShortened output from the actual CLI:
PIIFlow 0.3.0
PIIFLOW001 INFO Sanitized email data reaches an external LLM.
Source: redacted_email.py:6 request.json['email']
Path: request.json['email'] -> email -> redact_email(email)
Sink: redacted_email.py:7 client.responses.create(input=)
Policy: compliant action=require_sanitizer rule=require-redacted-email-openai
Provider: openai
Trust zone: external
Required sanitizers: redact_email
Observed sanitizers: redact_email
Reason: Email may be sent to OpenAI only after redact_email.
flowchart LR
A[Python source] --> B[AST and indexing]
B --> C[Structured taint propagation]
C --> D[Sensitive flow]
D --> E{Policy supplied?}
E -->|No| F[Text or JSON report]
E -->|Yes| G[Policy evaluator]
G --> H[Compliant record or violation]
H --> F
Policy evaluation is layered on top of flow analysis. The analyzer records a sensitive flow with category, source, path, sink, sanitizer evidence, provider, and trust zone. The policy evaluator annotates that flow without changing the flow fingerprint.
See docs/architecture.md and docs/policy-model.md.
PIIFlow requires Python 3.11 or newer. From a repository checkout:
uv sync --extra devBuild and install the package locally:
uv build
uv pip install dist/piiflow-0.3.0-py3-none-any.whlThe package is not claimed to be published on PyPI.
uv run piiflow scan examples/safe_app
uv run piiflow scan examples/vulnerable_app
uv run piiflow scan examples/vulnerable_app --format json
uv run piiflow scan examples/policy_apps/redacted_email.py --policy examples/policies/sanitized_email.yml
uv run python benchmarks/run.py --partition all --system all --limit 6The vulnerable scan exits 1 because findings meet the default high threshold. The safe and compliant policy examples exit 0.
piiflow scan PATH [OPTIONS]Options:
Exit codes:
Packaged defaults live in src/piiflow/rules/default.yaml. Custom YAML extends the defaults:
version: 1
sensitive_names:
CUSTOMER_SECRET:
- loyalty_id
sink_patterns:
- id: custom-agent
call: "*.invoke"
arguments: [0, "input"]
destination: configurable_llm
provider: generic
trust_zone: external
sanitizer_patterns:
- redact_piiHashing and encoding are ordinary transformations unless configured as sanitizers.
Policy schema version 1 supports ordered first-match rules:
version: 1
default_action: deny
rules:
- id: allow-local-financial
description: Financial data may be sent to local models.
categories: [FINANCIAL]
providers: [local]
trust_zones: [local]
action: allowFields:
Built-in sink metadata:
| Sink family | Provider | Trust zone |
|---|---|---|
| OpenAI Responses and Chat Completions | openai | external |
| Anthropic Messages | anthropic | external |
| Generic *.invoke sinks | generic | external |
Custom sink rules may override provider and trust_zone. The examples also show local/local and generic/internal configurations.
The benchmarks are synthetic. They measure scoped behavior in small Python snippets and should not be read as proof of real-world generalization, legal compliance, or superiority over other analyzers.
Original frozen 90-case benchmark:
| System | Cases | Precision | Recall | F1 | Accuracy | TP | FP | TN | FN |
|---|---|---|---|---|---|---|---|---|---|
| lexical | 90 | 0.714 | 1.000 | 0.833 | 0.778 | 50 | 20 | 20 | 0 |
| direct | 90 | 1.000 | 0.060 | 0.113 | 0.478 | 3 | 0 | 40 | 47 |
| piiflow_legacy | 90 | 0.909 | 0.800 | 0.851 | 0.844 | 40 | 4 | 36 | 10 |
| piiflow | 90 | 0.955 | 0.840 | 0.894 | 0.889 | 42 | 2 | 38 | 8 |
Frozen original benchmark hashes:
Phase 2B1 structured-analysis ablation:
| System | Precision | Recall | F1 | Accuracy | Total runtime |
|---|---|---|---|---|---|
| v0.1.0 baseline | 0.909 | 0.800 | 0.851 | 0.844 | 0.382s |
| v0.2.0 structured analysis | 0.955 | 0.840 | 0.894 | 0.889 | 0.252s |
| Delta | +0.045 | +0.040 | +0.043 | +0.044 | -0.131s |
Policy benchmark:
| System | Cases | Precision | Recall | F1 | Accuracy | TP | FP | TN | FN |
|---|---|---|---|---|---|---|---|---|---|
| policy_blind | 36 | 0.556 | 1.000 | 0.714 | 0.556 | 20 | 16 | 0 | 0 |
| piiflow_policy | 36 | 1.000 | 1.000 | 1.000 | 1.000 | 20 | 0 | 16 | 0 |
PIIFlow achieved 1.00 F1 on the 36-case synthetic policy benchmark. This means only that the current policy evaluator matched those scoped synthetic labels.
Policy benchmark hashes:
Full reports:
uv sync --extra dev
uv run ruff format --check .
uv run ruff check .
uv run mypy src/piiflow benchmarks
uv run pytest --cov=piiflow --cov-report=term-missing --cov-fail-under=90
uv run python benchmarks/run.py --partition all --system all
uv run python benchmarks/run_policy.py --system all
uv buildSee docs/reproducibility.md for environment setup, hash checks, wheel smoke testing, expected exit codes, and timing caveats.
PIIFlow supports common explicit data-flow patterns: request subscripts and getters, direct and annotated assignments, simple containers, known dictionary fields, known object attributes, f-strings, string concatenation, conditionals, awaited calls, local helper calls, simple imports, instance/static/class method binding, sanitizer calls, and whole-object sink arguments.
Current known benchmark misses are concentrated in closures, nested functions, dynamic attributes, dynamic sink construction, exception-flow propagation, framework-generated values, generators, and implicit/control-flow taint.
Current known false positives include decorator-provided sanitization and runtime target checks that are intentionally outside the analyzer's static model.
See docs/limitations.md for the complete threat-model and limitations discussion.
PIIFlow scans source code locally and does not execute target applications. It does not call LLM providers, send code to a remote service, or validate legal privacy compliance. Do not put genuine personal data, credentials, private keys, or proprietary code snippets in public issues or benchmark contributions.
Security reporting guidance is in SECURITY.md.
Contributions should keep benchmark labels honest, preserve reproducibility, and include tests for behavior changes. Start with CONTRIBUTING.md.
Citation metadata is available in CITATION.cff.
Apache-2.0. See LICENSE.
| Back | FazBrowse Home | New Git URL |