This dataset consists of chunked text extracted from all publicly available PDF documents on the Amazon Web Services (AWS) official website. The data includes user guides, technical documentation, and best practices about every AWS service, concept, and architecture.
It is designed to use in embedding generation, vector databases, and retrieval-augmented generation (RAG) systems.
The dataset is provided as a .json file.⦠See the full description on the dataset page:
https://huggingface.co/datasets/semihk1/aws-public-pdf-chunked-dataset.