README.md

former

Simple transformer implementation from scratch in pytorch. See http://peterbloem.nl/blog/transformers for an in-depth explanation.

Limitations

The current models are designed to show the simplicity of transformer models and self-attention. As such they will not scale as far as the bigger transformers. For that you'll need a number of tricks that complicate the code (see the blog post for details).

All models so far are a single stack of transformer blocks (that is, no encoder/decoder structures). It turns out that this simple configuration often works best.

Use

You can clone the code and run the experiments from the root directory. E.g.

python experiments/classify.py

Hyperparameters are passed as command line arguments. The defaults should work well. The classification data is automatically downloaded, the wikipedia data is included in the repository.

You should be able to install as a package as well, with

pip install git+https://github.com/pbloem/former

but I haven't tried this. It's probably easier to just copy over the code you need. Let me know if you need this for anything and it doesn't work.

Requirements

Python 3.6+ is required.

The following should install all requirements pip install torch tb-nightly tqdm numpy torchtext

You may also need pip install future depending on the exact python version.

GitHub - pbloem/former: Simple transformer implementation from scratch in pytorc...

README.md

former

Limitations

Use

Requirements

Recommend

十步法原则解决数据质量问题 - 宜信技术学院

深度长文：细说iOS代码签名 - xelz's blog

Sneak Peek: First Chapter From New SwiftUI, Combine, and Catalyst Books! [FREE]

GitHub - GoogleCloudPlatform/cloud-run-button: Let anyone deploy your GitHub rep...

GitHub - asaskevich/govalidator: [Go] Package of validators and sanitizers for s...

A simple bookmarklet to tweet the current page

Nginx+Zuul集群实现高可用网关

Is This Magic!? Ferris Explores rustc

C preprocessor tricks, tips, and idioms (2015)

The difference between isolation levels and consistency levels

About Joyk