Customize OpenAI privacy-filter for Snowflake semantic_categories
ActiveFine-tune OpenAI's privacy-filter model on Snowflake's semantic_categories taxonomy, then evaluate against a hand-labeled holdout.
Overview
OpenAI’s privacy-filter model is trained on generic PII categories. Snowflake’s semantic_categories use a richer, domain-specific taxonomy (e.g. US_PASSPORT, IBAN_CODE, HEALTHCARE_NUMBER).
Goal: Customize the privacy-filter so it can output Snowflake-aligned categories directly, then evaluate quality against a hand-labeled holdout set.
Success criteria: (revised 2026-09-22 — see the entry below)
- Coverage ≥ 0.70 at a 1% wrong-tag budget: the share of columns auto-tagged with
SEMANTIC_CATEGORYwhile holding the wrong-tag rate under 1%, with the remainder backed off toPRIVACY_CATEGORYor escalated. - Macro-F1 ≥ 0.80 on the holdout across all Snowflake categories that have ≥ 20 labeled examples — retained as a secondary check, no longer the gate.
- Expected calibration error ≤ 0.05 out of domain, with temperature fitted on in-distribution dev rows only.
- Latency p95 ≤ 150 ms per document at evaluation time.
Resources
- Introducing OpenAI Privacy Filter — announcement with architecture details (1.5B params, 50M active, 128k context)
- openai/privacy-filter on GitHub — Python CLI (
opf) for redaction, evaluation, and fine-tuning - openai/privacy-filter on Hugging Face — model weights (Apache 2.0),
AutoModelForTokenClassificationusage - Community fine-tuned variants — existing fine-tunes including quantized and domain-specific adaptations
- Snowflake
EXTRACT_SEMANTIC_CATEGORIES— reference taxonomy (47 categories as of 8.x) - ai4privacy/pii-masking-300k — OpenPII-220k (27 PII classes, 6 languages) + FinPII-80k (~20 finance/insurance classes); ~98.3% label accuracy
- Snowflake ML Jobs overview — run arbitrary Python ML workloads on Snowflake GPU compute pools via
@remotedecorator orsubmit_*APIs - Snowflake distributed training —
PyTorchDistributorfor multi-node/multi-GPU training inside Container Runtime - SPCS AWS instance families — GPU_NV_S (1x A10G 24 GB) through GPU_NV_L (8x A100 320 GB) and GPU_L40S/GPU_R6K
- jaredpalmer/kev on GitHub — open System One decision model family (Kev-4B) on Qwen3.5 with TypeSafe API contract
- Jev on Snowflake plan — capability probe, Modal H100 reproduction, and Model Registry staging runbook
2026-05-06 — kickoff
- Pulled the full Snowflake
semantic_categoriesreference list (47 categories as of Snowflake 8.x). - Mapped each Snowflake category to the closest OpenAI
privacy-filteroutput label where one exists. ~30% have no direct mapping — these are the interesting gap cases. - Next: sample the gap categories from internal datasets and build a label schema for annotation.
2026-04-23 — initial exploration
Fine-tuning pipeline
opf trainships natively — handles JSONL ingestion, 128-token banded attention windowing, AdamW with gradient accumulation, checkpoint serialization.- Two demo workflows: policy adaptation (relabel existing categories) and new taxonomy (custom label space).
- Output head remapping: exact-match labels copy weights directly; new labels warm-start from closest base class (e.g.
B-custom_idinheritsB-IDweights). This is why OpenAI’s benchmark jumped from 54% → 96% F1 on small data. - Custom taxonomies configured via
label_space.jsonwithspan_class_names. BIOES expansion is automatic; background classOmust be first entry.
Compute requirements
- Model: 2.8 GB safetensors (BF16), 1.5B total params / 50M active (MoE).
- Full fine-tune: ~24 GB VRAM (mixed precision) — single A100 40GB or RTX 4090 24GB.
- LoRA alternative: ~12 GB, viable on consumer GPUs.
PoC plan (one-week target)
- Environment + baseline — clone repo, pull checkpoint (~17 GB), verify
opf redacton sample text, download ai4privacy English splits, convert toopfeval JSONL. - Baseline eval — collapse ai4privacy’s 27 classes → Privacy Filter’s 8 categories, run
opf eval, reproduce ~96% F1 as sanity check, note per-category breakdown. - Taxonomy design — design ~15–20 target categories aligned to Snowflake’s
SEMANTIC_CATEGORY(NAME, EMAIL, PAYMENT_CARD, PASSPORT, NATIONAL_IDENTIFIER, STREET_ADDRESS, PHONE_NUMBER, IP_ADDRESS, DATE_OF_BIRTH, AGE, GENDER, OCCUPATION, SALARY, MEDICAL_CONDITION, MEDICATION…). Writelabel_space.json, maintain a mapping file. - Data prep + smoke training — generate JSONL with new labels, 80/10/10 split, run
opf trainon 5–10k subset to validate pipeline end-to-end. - Full fine-tune — ~150k examples, 2–3 epochs, single A100, checkpoint every N steps.
- Evaluation —
opf evalon held-out test, build confusion matrix. Focus on overlapping pairs: NAME vs ORGANIZATION_IDENTIFIER, NATIONAL_IDENTIFIER vs TAX_IDENTIFIER, PAYMENT_CARD vs BANK_ACCOUNT.
Key risks
- Category overlap degradation — expect 10–20 pt F1 drop on overlapping numeric-ID categories.
- Label skew — ai4privacy over-represents
firstname; eval needs stratified sampling to avoid misleading macro-F1. - OOD gap — ai4privacy covers education/health/psychology/finance; Snowflake customer data (log lines, transaction records, support tickets) may differ.
- Licensing — ai4privacy is academic-friendly; production use at Snowflake needs commercial license or synthetic dataset alternative.
2026-07-20 — can we train this on Snowflake?
The question
The original plan assumed a local A100 or RTX 4090. The question is whether we can run the full opf train pipeline on the Snowflake platform instead — keeping training data, labeled holdout, and the fine-tuned checkpoint inside the Snowflake trust boundary. The short answer is yes, via Snowflake ML Jobs on a GPU compute pool, but the product fit is narrow. The two turnkey Cortex fine-tuning offerings both miss by model family, so the right path is the general-purpose container runtime.
Paths considered
Cortex Fine-Tuning (GA) — LoRA-only, six pre-approved generative base models (llama3-8b/70b, llama3.1-8b/70b, mistral-7b, mixtral-8x7b), strict prompt/completion schema. Not applicable: privacy-filter is an AutoModelForTokenClassification, not one of those models.
Cortex Training (Public Preview, June 2026) — full fine-tune + reinforcement learning via managed GPU pools and ArcticTraining YAML config. Supported families: open-weight Qwen and Mistral only. Again, not applicable for this model.
Snowflake ML Jobs + Container Runtime on a GPU compute pool — lift-and-shift of the existing opf train / HF transformers loop. Custom pip and HuggingFace packages install via runtime_environment. The @remote decorator or submit_from_stage submits the job; PyTorchDistributor handles multi-GPU data-parallel scaling. This is the right path.
Model Registry + SPCS serving — after training, the fine-tuned checkpoint registers as a token-classification HF pipeline (TransformersPipeline) and deploys via create_service on a GPU pool. The holdout eval and p95 latency check run entirely in-account.
Revised plan (Snowflake ML Jobs)
Replaces the local A100/4090 assumption from the 2026-04-23 PoC plan. Full setup walkthrough in docs/plans/privacy-filter/snowflake-setup.md. For a lower-friction serverless PoC alternative, see docs/plans/privacy-filter/modal-setup.md.
- Data staging — load ai4privacy English splits and any hand-labeled Snowflake-category examples into a Snowflake internal stage or table. Convert to
opfJSONL inside a Container Runtime job using a CPU pool. - Smoke training — submit
opf train(5–10k examples, 1 epoch) as an ML Job onGPU_NV_S(1× A10G, 24 GB VRAM). This covers the LoRA path (~12 GB) and is the cheapest smoke check; validates the end-to-end pipeline without burning large-GPU credits. - Full fine-tune — scale to
GPU_L40S(up to 8× L40S, 48 GB each) orGPU_NV_L(8× A100, 40 GB each, on request) for the ~150k-example, 2–3 epoch run. UsePyTorchDistributorwithnum_nodes=1, num_gpus=4as the baseline; add nodes if epoch time exceeds a few hours. - Checkpoint to stage — write final checkpoint safetensors to a named Snowflake stage. Pin the commit hash and epoch number in the stage path for reproducibility.
- Register + serve — log the fine-tuned model to the Model Registry from the stage (custom model path, not the HF repo id — see trade-offs below). Call
create_serviceon a GPU pool for the holdout eval endpoint. Runopf evalagainst the held-out test set in-account; build the confusion matrix from the service response.
Trade-offs
GPU pool availability — GPU_NV_L (A100) requires an on-request allocation; not always available in all AWS regions. GPU_L40S (L40S, 48 GB/GPU) and GPU_R6K (Blackwell RTX 6000, 96 GB/GPU) are GA as of May 2026 but AWS-only. Plan for GPU_NV_M (4× A10G, 96 GB total) as the fallback if neither is immediately available — it’s sufficient for a distributed full fine-tune.
Data residency win — this is the primary reason to prefer Snowflake over local: training data never leaves the account boundary. For a model intended to run on Snowflake customer data (log lines, transaction records, support tickets), keeping the labeled examples and the training run in the same governance perimeter is meaningful.
Model Registry caveat — the TransformersPipeline wrapper in the Model Registry takes an HF Hub repo identifier by default. A custom fine-tuned checkpoint on a Snowflake stage needs to be logged either (a) via a custom snowflake.ml.model.custom_model.CustomModel wrapper that loads from the stage path, or (b) by pushing the checkpoint to a private HF repo first and referencing it. Option (a) keeps everything in-account; option (b) is simpler to implement. Decide before step 5.
2026-09-22 — coverage at an error budget replaces macro-F1
What changed
The gate moves from macro-F1 to coverage at an error budget, and calibration becomes a tracked metric rather than an afterthought. Reasoning is in Decision models: accuracy converges, calibration does not; the project-specific part is here.
Macro-F1 scores the classifier as if it must name one of the 47 categories every time. That throws away the thing the model already knows — which of its own predictions it should not have made. Nothing downstream wants a forced guess. A wrong SEMANTIC_CATEGORY drives a wrong tag, which drives a wrong masking policy.
The backoff Snowflake already supports
Snowflake classification assigns every column two tags: SNOWFLAKE.CORE.SEMANTIC_CATEGORY (fine, 47 values) and SNOWFLAKE.CORE.PRIVACY_CATEGORY (coarse — IDENTIFIER, QUASI_IDENTIFIER, SENSITIVE). See Introduction to sensitive data classification. The hierarchy is already there, which makes the coarse answer free: it follows from the fine prediction with no second call.
So the serving contract becomes three-way instead of one-way:
| Confidence | Emit | Consumer |
|---|---|---|
| Above the fine threshold | SEMANTIC_CATEGORY | auto-tag |
| Between | PRIVACY_CATEGORY only | auto-tag, coarse |
| Below | nothing | governance review queue |
TypeSafe’s SEC filings cookbook is the worked version of this on a different taxonomy: 75 SIC groups over 60 filings, 39/60 forced to name a group, 48/60 useful answers once unsure cases report the parent division instead. The confident half was right 27/30; the unsure half 12/30, which becomes 70% one level up.
Evaluation protocol, borrowed from Kev
Jared Palmer’s Kev is methodologically the closest published work to this project — a LoRA plus pointer head on a small open base, trained for a closed decision space, evaluated against a commercial decision model. The protocol is more rigorous than anything in the privacy-filter material and is worth copying wholesale:
- Locked test partition, read once per candidate. Development partitions select the model. The test partition is read at most once and the read decides promotion. Every current number in this project comes from a partition that would not survive that rule.
- Temperature fitted on in-distribution dev rows only, baked into the head so every loader gets calibrated probabilities by default. It never changes an answer, so accuracy is identical in both columns.
- Report raw and calibrated side by side. Temperature scaling is monotone: it repairs average miscalibration and cannot reorder confidences. Kev’s ECE beats Jev’s (0.041 vs 0.049) while its coverage does not (0.57 vs 0.70). Reporting ECE alone would have hidden that.
- Suite hash, code hash, and git commit in every result file, with paired bootstraps on every comparison.
- A deliberately unknowable slice. Kev measures the share of unanswerable items answered at confidence ≥ 0.9 (Kev 0.00, Jev 0.09). For PII this is the column whose name and samples genuinely do not determine a category, and it is the failure that matters most.
Revised evaluation steps
Replaces step 6 of the 2026-07-20 plan.
- Split the hand-labeled holdout into dev and a locked test partition before any tuning. Record hashes.
- Run
opf evalon dev for the confusion matrix, still focusing on the overlapping pairs (NAME vs ORGANIZATION_IDENTIFIER, NATIONAL_IDENTIFIER vs TAX_IDENTIFIER, PAYMENT_CARD vs BANK_ACCOUNT). - Fit temperature on dev in-distribution rows, minimizing NLL. Store it with the checkpoint.
- Sweep the two thresholds against the 1% wrong-tag budget to produce the coverage number.
- Build the unknowable slice from columns whose name and samples are genuinely ambiguous; measure the share answered above the fine threshold.
- Read the locked test once.
Open questions
- Does a span-level token classifier give a usable confidence at all? Coverage assumes one scalar per decision. A BIOES tagger emits a distribution per token, and column-level classification aggregates over sampled values. The aggregation rule is undecided and it determines whether any of this is measurable.
- What is the actual wrong-tag cost? 1% is a guess. The right budget comes from what a wrong masking policy costs relative to a column sitting in the review queue, and I have not priced either.
- No rationale. A confidence-gated tag tells an auditor how sure the model was, not what it fired on. A token classifier can at least return the span. Worth keeping that output even where only the column-level label is consumed.