Comments (2)
Thanks for looking into this! It turns out terminal was running an old version of java http://stackoverflow.com/questions/12757558/installed-java-7-on-mac-os-x-but-terminal-is-still-using-version-6
from tabula-py.
I can't reproduce your error.
~/t/tabula-py (master ☡=) (default) ipython 15:41:58
Python 3.5.2 (default, Oct 11 2016, 05:05:28)
Type "copyright", "credits" or "license" for more information.
IPython 5.1.0 -- An enhanced Interactive Python.
? -> Introduction and overview of IPython's features.
%quickref -> Quick reference.
help -> Python's own help system.
object? -> Details about 'object', use 'object??' for extra details.
In [1]: import tabula
In [2]: df = tabula.read_pdf("https://github.com/tabulapdf/tabula-java/raw/master/src/test/resources/technology/tabula/arabic.pdf")
Picked up JAVA_TOOL_OPTIONS: -Dfile.encoding=UTF-8
In [3]: df
Out[3]:
مرحبًا اسمي سلطان
0 انا من ولاية كارولينا الشمال من اين انت؟
1 1234 عندي 47 قطط
2 هل انت شباك؟ اسمي Jeremy في الانجليزية
3 Jeremy is جرمي in Arabic NaN
In [4]: print(df)
مرحبًا اسمي سلطان
0 انا من ولاية كارولينا الشمال من اين انت؟
1 1234 عندي 47 قطط
2 هل انت شباك؟ اسمي Jeremy في الانجليزية
3 Jeremy is جرمي in Arabic NaN
It seems tabula-java issues. How about using tabula-java directly as follows:
java -jar /path/to/jar/tabula-0.9.1-jar-with-dependencies.jar -p 1 -g arabic.pdf
from tabula-py.
Related Issues (20)
- Allow columns parameter to use relative area HOT 5
- Cutting off first character of last column HOT 5
- tabula.io.read_pdf 'columns' argument change typing to Iterable[float] HOT 1
- tabula.io.read_pdf 'columns' argument change type annotation to Iterable[float] HOT 3
- tabula.io.read_pdf argument "pandas_options" is being changed inside the function HOT 1
- tabula.io.read_pdf argument "pandas_options" is being changed inside the function HOT 3
- Extracting non tabular data from pdfs, is it possible? HOT 1
- Extracting non-tabular (1-tabula output) data from pdf, is it possible? HOT 3
- Unable to remove error : Got stderr: Picked up _JAVA_OPTIONS: -Djava.awt.headless=true HOT 1
- Unable to remove note in log : Got stderr: Picked up _JAVA_OPTIONS: -Djava.awt.headless=true HOT 1
- Tabula py Ignores an entire column if it's blank and if it does not contain headerd? HOT 1
- tabula-py CalledProcessError: Command '['java', '-Dfile.encoding=UTF8', '-jar', HOT 3
- dont ignore empty columns in tables spanning multiple pages HOT 1
- Try to install tabula-py HOT 1
- Use JPype instead of subprocess HOT 11
- Add a way to set areas for non-existent pages in template HOT 4
- Exception: RuntimeError: java.lang.UnsatisfiedLinkError: HOT 2
- cant install tabula-py on m1 mac vscode. HOT 1
- Support Python 3.12 HOT 5
- Pls add "orientation" parameter to read_pdf HOT 4
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from tabula-py.