> For the complete documentation index, see [llms.txt](https://aws-gcr-wwso-security.gitbook.io/an-quan-zui-jia-shi-jian/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://aws-gcr-wwso-security.gitbook.io/an-quan-zui-jia-shi-jian/3.-shu-ju-bao-hu/shi-bie-ji-biao-shi-min-gan-shu-ju/min-gan-shu-ju-sao-miao-yu-ni-ming-hua-kai-yuan-gong-ju.md).

# 敏感数据扫描与匿名化开源工具

{% tabs %}
{% tab title="敏感数据扫描" %}
Presidio Analyzer - [地址](https://microsoft.github.io/presidio/analyzer/)

The Presidio analyzer is a Python based service for detecting PII entities in text.

During analysis, it runs a set of different *PII Recognizers*, each one in charge of detecting one or more PII entities using different mechanisms.

Presidio analyzer comes with a set of predefined recognizers, but can easily be extended with other types of custom recognizers. Predefined and custom recognizers leverage regex, Named Entity Recognition and other types of logic to detect PII in unstructured text.

PrivAPI-[地址](https://github.com/bballamudi/privapi)

*PrivAPI* is a Python package that allows the classification of sensitive data flows within REST API communication using Deep Neural Networks (DNN). It relies on Google's Keras and TensorFlow.

d18n - [地址](https://github.com/LianjiaTech/d18n/blob/main/doc/detect.md)

d18n has a built-in method, use keywords match, regexp match, and NLP words match.

Octopii - [地址](https://github.com/redhuntlabs/octopii)

Octopii is an open-source AI-powered Personal Identifiable Information (PII) scanner that can look for image assets such as Government IDs, passports, photos and signatures in a directory.

[piicatcher](https://github.com/tokern/piicatcher) - [地址](https://github.com/tokern/piicatcher)

PIICatcher is a scanner for PII and PHI information. It finds PII data in your databases and file systems and tracks critical data.
{% endtab %}

{% tab title="匿名化工具" %}
Presidio Anonymizer - [地址](https://microsoft.github.io/presidio/anonymizer/)

The Presidio anonymizer is a Python based module for anonymizing detected PII text entities with desired values. Presidio anonymizer supports both anonymization and deanonymization by applying different operators. Operators are built-in text manipulation classes which can be easily extended.

Kafka Connect Pipeline - [地址](https://github.com/Stefen-Taime/Kafka-pipeline)

Learn how to build a data pipeline using a combination of open-source software (OSS), including Debezium, Apache Kafka, Kafka Connect.

d18n - [地址](https://github.com/LianjiaTech/d18n/blob/main/doc/mask.md)

Data mask
{% endtab %}
{% endtabs %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation by asking a question.

Perform an HTTP GET request on the following URL with the `ask` and `goal` query parameters:

```
GET https://aws-gcr-wwso-security.gitbook.io/an-quan-zui-jia-shi-jian/3.-shu-ju-bao-hu/shi-bie-ji-biao-shi-min-gan-shu-ju/min-gan-shu-ju-sao-miao-yu-ni-ming-hua-kai-yuan-gong-ju.md?ask=<question>&goal=<user_goal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is what the user is ultimately trying to achieve, the reason they need the answer. Sharing it helps GitBook give you a better, more relevant answer. A goal is most helpful when it describes the outcome the user wants rather than restating the question. For example, with `ask=how do I create an API token`, a goal like `automate deployments from our CI pipeline` lets GitBook tailor the answer to that use case.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
