What is actually stored in a PDF’s metadata?

Two places hold it. The document information dictionary is the old set of fields: title, author, subject, keywords, creator (the program you typed in), producer (the library that wrote the file), creation date and modification date. The XMP packet is an XML block that often repeats those and adds more, including editing history from design software.

Neither shows on the page. Both travel with the file, survive email, and are read by any reader, search index or forensic tool that opens it.

Why does it matter who made the file?

An author field is a name, and a name is a person. Anonymous submissions, whistleblower documents, job applications sent through an agency, tender responses and legal filings have all been traced back through this field.

Dates matter as much. A "new" contract with a creation date from three years ago, or a modification date two minutes before it was sent, tells a story the sender did not intend to tell.

What metadata cannot be removed?

Text and images on the page are content, not metadata: if a name is printed on the page, redact it instead. Images inside the PDF may carry their own EXIF, including GPS coordinates from a phone camera; stripping the PDF’s fields does not touch those, so take the pictures out and check them if the source is a phone.

Digital signatures also embed the signer’s details by design — removing them would defeat the signature.

Before sending anything sensitive, run the privacy scanner over the finished file: it reports metadata, hidden text, attachments and scripts together.