For academic reference, cite the following paper:
https://ieeexplore.ieee.org/document/10223689
CryptoBERT is a pre-trained NLP model to analyse the language and sentiments of cryptocurrency-related social media posts and messages. It was built by further training the
vinai's bertweet-base language model on the cryptocurrency domain, using a corpus of over 3.2M unique cryptocurrency-related social media posts.
(A research paper with more details will follow soon.)
The model was trained on the following labels: "Bearish" : 0, "Neutral": 1, "Bullish": 2
CryptoBERT's sentiment classification head was fine-tuned on a balanced dataset of 2M labelled StockTwits posts, sampled from
ElKulako/stocktwits-crypto.
CryptoBERT was trained with a max sequence length of 128. Technically, it can handle sequences of up to 514 tokens, however, going beyond 128 is not recommended.
1from transformers import TextClassificationPipeline, AutoModelForSequenceClassification, AutoTokenizer
2model_name = "ElKulako/cryptobert"
3tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=True)
4model = AutoModelForSequenceClassification.from_pretrained(model_name, num_labels = 3)
5pipe = TextClassificationPipeline(model=model, tokenizer=tokenizer, max_length=64, truncation=True, padding = 'max_length')
6# post_1 & post_3 = bullish, post_2 = bearish
7post_1 = " see y'all tomorrow and can't wait to see ada in the morning, i wonder what price it is going to be at. 😎🐂🤠💯😴, bitcoin is looking good go for it and flash by that 45k. "
8post_2 = " alright racers, it’s a race to the bottom! good luck today and remember there are no losers (minus those who invested in currency nobody really uses) take your marks... are you ready? go!!"
9post_3 = " i'm never selling. the whole market can bottom out. i'll continue to hold this dumpster fire until the day i die if i need to."
10df_posts = [post_1, post_2, post_3]
11preds = pipe(df_posts)
12print(preds)
13
14
CryptoBERT was trained on 3.2M social media posts regarding various cryptocurrencies. Only non-duplicate posts of length above 4 words were considered. The following communities were used as sources for our corpora:
(1) StockTwits - 1.875M posts about the top 100 cryptos by trading volume. Posts were collected from the 1st of November 2021 to the 16th of June 2022.
ElKulako/stocktwits-crypto
(2) Telegram - 664K posts from top 5 telegram groups:
Binance,
Bittrex,
huobi global,
Kucoin,
OKEx.
Data from 16.11.2020 to 30.01.2021. Courtesy of
Anton.
(3) Reddit - 172K comments from various crypto investing threads, collected from May 2021 to May 2022
(4) Twitter - 496K posts with hashtags XBT, Bitcoin or BTC. Collected for May 2018. Courtesy of
Paul.