This repository is home to the Falcon-40B-Instruct model, which has been carefully converted from its original 32-bit mode to an efficient and compact 8-bit file.
You can use this model directly with a pipeline for tasks such as text generation and instruction following:
1from transformers import pipeline
2
3generator = pipeline('text-generation', model='tensorcat/falcon-40b-instruct-8bit')
4print(generator("Generate a story about a spaceship traveling through space.", max_length=200))
5