Comments (7)
we have previously used jieba for word segmentation, then used BPE to learn a subword segmentation: http://www.statmt.org/wmt17/pdf/WMT39.pdf
if you want to try word segmentation with BPE, you can run learn_bpe.py on unsegmented text, but note that the code was not optimized for this use case, and will likely be slower and consume more memory.
from subword-nmt.
@parahaoer for Chinese word segmentation, you can use this tool jieba or stanford parser
from subword-nmt.
@liesun1994 Could you please tell me how you do it?
from subword-nmt.
THX
from subword-nmt.
@rsennrich I used a transformer base model in fairseq as specified in paper with same no of merge (59500 - separate BPE models) operations that gave around 24 M sentences except back translated data .
But I am getting low BLEU score of 17.69 (on 2017 test data) as compared to what your paper claims I use sacre BLEU to compute BLEU score which gives same BLEU as official MT eval script
I could not find out the reason to atleast achieve the baseline
from subword-nmt.
this question/complaint is very underspecified (you don't even mention the translation direction), but I don't expect your problem to be the BPE segmentation.
You can actually find most of our WMT17 submissions here: http://data.statmt.org/wmt17_systems/
Maybe this will help you find the difference to your training setup, for example in preprocessing.
from subword-nmt.
The translation direction is en -> zh
The problem is in evaluating BLEU as I haven't properly used tokenizeChinese.py as specified. Got 32.7
Since the doubt pertains to chinese data I have posted in the same issue rather than creating a new one. Apologies
from subword-nmt.
Related Issues (20)
- subword-nmt HOT 3
- learn_bpe.py error HOT 1
- learn_joint_bpe_and_vocab.py for Japanese HOT 1
- BPE-Dropout question HOT 1
- Recover back code file HOT 1
- DeprecationWarning and ResourceWarning: Enable tracemalloc to get the object allocation traceback HOT 2
- No module named apply_bpe HOT 4
- About the vocabulary size HOT 2
- Readme update please HOT 1
- Question about vocabulary filter HOT 2
- Unknown word and vocabulary filter HOT 2
- How to avoid special char like '\t' being split by bpe HOT 3
- applying BPE(Byte Pair Encoding) fails for large Chinese data tokenized with THULAC HOT 1
- Error: invalid line 2 in BPE codes file when running apply_bpe.py HOT 1
- How to decode BPE when applied to machine translation HOT 1
- learn_bpe.py code question HOT 1
- BrokenPipeError: [Errno 32] Broken pipe HOT 3
- Is it possibile to extend a trained BPE model's merge operations? HOT 2
- How to find all valid BPEs for a word? HOT 3
- AttributeError: 'BPE' object has no attribute 'glossaries_regex' HOT 1
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from subword-nmt.