페르소나 기반 키워드 top1 선택 데이터셋. 350m 페르소나 모델의 <|kw|> 태스크 학습용.
(persona, date, 2h block) 마다: 그 블록에서 총 15분 이상 시청한 프로그램들의 키워드(1~8, dedup) 중,
그 사람의 성향을 가장 잘 대표하는 **키워드 1개(top1)**를 고른다.
입력: 주간 페르소나(전주, key_keywords 제외: viewing_tendency/preferred_genres/preferred_programs/frequent_channels) + 직전일 일일페르소나 서술 + 후보 키워드(1~8). 최근 시청이력은 입력에 넣지 않음(후보를 만들 때만 사용).
출력: 키워드 1개 (label_keyword)
후보 키워드 = 그 블록 15분+ 시청 프로그램들의 키워드(프로그램별 1개, dedup). 정답 = 교사… See the full description on the dataset page:
https://huggingface.co/datasets/HBKenerzai/LGUplus_keyword_reranking.