This dataset is released as part of the ICICLE project and is used for training models to classify organizations into Standard Industrial Classification (SIC) codes. The dataset contains multiple textual representations of organizational descriptions collected via web search and large language model (LLM) summarization.