Views
No views yet
firstnamemiddlenamelastnamesexdob (Date of Birth)agegenderheighteyecoloremailphonenumberurlusernameuseragentstreetcitystatecountyzipcodecountrysecondaryaddressbuildingnumberordinaldirectionnearbygpscoordinatecompanynamejobtitlejobareajobtypeaccountnameaccountnumbercreditcardnumbercreditcardcvvcreditcardissueribanbiccurrencycurrencynamecurrencysymbolcurrencycodeamountpinssnimei (Phone IMEI)mac (MAC Address)vehiclevin (Vehicle VIN)vehiclevrm (Vehicle VRM)bitcoinaddresslitecoinaddressethereumaddressip (IP Address)ipv4ipv6maskednumberpasswordtimeordinaldirectionprefix1### Instruction:
2 Identify and extract the following PII entities from the text, if present: companyname, pin, currencyname, email, phoneimei, litecoinaddress, currency, eyecolor, street, mac, state, time, vehiclevin, jobarea, date, bic, currencysymbol, currencycode, age, nearbygpscoordinate, amount, ssn, ethereumaddress, zipcode, buildingnumber, dob, firstname, middlename, ordinaldirection, jobtitle, bitcoinaddress, jobtype, phonenumber, height, password, ip, useragent, accountname, city, gender, secondaryaddress, iban, sex, prefix, ipv4, maskednumber, url, username, lastname, creditcardcvv, county, vehiclevrm, ipv6, creditcardissuer, accountnumber, creditcardnumber. Return the output in JSON format.
3
4### Input:
5 Greetings, Mason! Let's celebrate another year of wellness on 14/01/1977. Don't miss the event at 176,Apt. 388.
6
7### Output:
8transformers library installed:pip install transformers1from transformers import AutoTokenizer, AutoModelForTokenClassification
2
3# Load the tokenizer and model
4tokenizer = AutoTokenizer.from_pretrained("ab-ai/PII-Model-Phi3-Mini")
5model = AutoModelForTokenClassification.from_pretrained("ab-ai/PII-Model-Phi3-Mini")
6
7
8input_text = "Hi Abner, just a reminder that your next primary care appointment is on 23/03/1926. Please confirm by replying to this email Nathen15@hotmail.com."
9
10model_prompt = f"""### Instruction:
11 Identify and extract the following PII entities from the text, if present: companyname, pin, currencyname, email, phoneimei, litecoinaddress, currency, eyecolor, street, mac, state, time, vehiclevin, jobarea, date, bic, currencysymbol, currencycode, age, nearbygpscoordinate, amount, ssn, ethereumaddress, zipcode, buildingnumber, dob, firstname, middlename, ordinaldirection, jobtitle, bitcoinaddress, jobtype, phonenumber, height, password, ip, useragent, accountname, city, gender, secondaryaddress, iban, sex, prefix, ipv4, maskednumber, url, username, lastname, creditcardcvv, county, vehiclevrm, ipv6, creditcardissuer, accountnumber, creditcardnumber. Return the output in JSON format.
12
13 ### Input:
14 {input_text}
15
16 ### Output: """
17
18
19inputs = tokenizer(model_prompt, return_tensors="pt").to(device)
20# adjust max_new_tokens according to your need
21outputs = model.generate(**inputs, do_sample=True, max_new_tokens=120)
22response = tokenizer.decode(outputs[0], skip_special_tokens=True)
23print(response) #{'middlename': ['Abner'], 'dob': ['23/03/1926'], 'email': ['Nathen15@hotmail.com']}
24