hive, spark で使えそうなサンプルデータ
Lahman’s Baseball Database
http://seanlahman.com/baseball-archive/statistics/
The Fake Name Generator
http://www.fakenamegenerator.com/order.php
Sample Datasets from STAR Experiment
https://sdm.lbl.gov/fastbit/data/samples.html
hypertable
https://code.google.com/p/hypertable/downloads/detail?name=access.tsv.gz
2015年1月26日月曜日
2015年1月17日土曜日
scala @ centos7
これがいいのかわからないけど、とりあえずバイナリを利用する。
(たぶんこれが一番簡単だと思う)
https://sites.google.com/site/scalajp/home/installation
を参考に。
http://www.scala-lang.org/download/ から
scala-2.11.5.tgz をダウンロード。
# mv scala-2.11.5 /usr/local/share
# cd /usr/local/share
# ln -s scala-2.11.5 scala
/etc/profile.d/scala.sh を作成しておく
(自分の.bashrc に書いてもOK)
(たぶんこれが一番簡単だと思う)
https://sites.google.com/site/scalajp/home/installation
を参考に。
http://www.scala-lang.org/download/ から
scala-2.11.5.tgz をダウンロード。
# mv scala-2.11.5 /usr/local/share
# cd /usr/local/share
# ln -s scala-2.11.5 scala
/etc/profile.d/scala.sh を作成しておく
(自分の.bashrc に書いてもOK)
export SCALA_HOME=/usr/local/share/scala export PATH=$PATH:$SCALA_HOME/bin
# source /etc/profile.d/scala.sh
# scala -version
> Scala code runner version 2.11.5 -- Copyright 2002-2013, LAMP/EPFL
REPLで動作確認
# scala
Welcome to Scala version 2.11.5 (OpenJDK 64-Bit Server VM, Java 1.7.0_75).
Type in expressions to have them evaluated.
Type :help for more information.
scala> println("Hello World")
Hello World
次はmavenで。
2014年11月22日土曜日
Hiveのcollect_list
以前はプリミティブなタイプしかつかえなかったが、
0.13からstructやarrayのも使えるようになった。
https://issues.apache.org/jira/browse/HIVE-5294
0.13からstructやarrayのも使えるようになった。
https://issues.apache.org/jira/browse/HIVE-5294
2014年11月8日土曜日
Hiveのチューニング
少し古いけどオススメ
http://www.slideshare.net/ye.mikez/hive-tuning
2013年7月 (Hive 0.11頃?)
ORCフォーマット(非圧縮)、Tez、パーティション(とバーチャルカラム?)の利用
ソートされたデータを入れる
short circuit read (HDFS-2246)
その他プロパティーの確認項目あり
現在のプロパティーの確認方法
hive>set;
クエリーのパフォーマンス比較 (2014年2月)
https://amplab.cs.berkeley.edu/benchmark/
RedShift, Impala, Shark(SparkSQL), Hive, Tez を比較
(HiveのORCFile, ImpalaのParquetは利用してない)
これを見るとMapReduceからTezに変更するだけで3~4割は早くなる感じ。
すごいのはSparkSQL。Hiveとどの程度互換性があるのか早く確認したい。
Tezも既にHiveに取り込まれて、簡単に利用できるので魅力的。
■ Tezの利用は簡単で効果大。とりあえずTezを利用すべし。
個人的な関心はSparkSQLへ。
Hive on Tez
https://cwiki.apache.org/confluence/display/Hive/Hive+on+Tez
0.13からにHiveに組み込まれている
https://issues.apache.org/jira/browse/HIVE-4660
https://issues.apache.org/jira/browse/HIVE-6098
Tez (Hive Configuration Properties)
https://cwiki.apache.org/confluence/display/Hive/Configuration+Properties#ConfigurationProperties-Tez
Chapter 10. Installing and Configuring Apache Tez
http://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.1.5/bk_installing_manually_book/content/rpm-chap-tez.html
■ tezを有効にする(mapredueからtezへ)
set hive.execution.engine=tez;
mapreduceへ戻す
set hive.execution.engine=mr;
2014年10月30日木曜日
Hue Hue だよ
Hueはかっこいい。しかもDjango(ver1.4だけど)でできてる。
(tomcatじゃないよ!!!)
画面はBootstrap, Knockout.js, jQuery
GitHub
https://github.com/cloudera/hue
Hue
http://gethue.com/
Hadoopの標準GUI HUEの最新情報
http://www.slideshare.net/Cloudera_jp/hadoopgui-hue
Hue SDK
http://cloudera.github.io/hue/docs-3.6.0/sdk/sdk.html
Hue関連のニュース
http://gethue-jp.tumblr.com/
インストール、設定
Install Hue without Cloudera
http://stackoverflow.com/questions/20579357/install-hue-without-cloudera
config
http://docs.hortonworks.com/HDPDocuments/HDP1/HDP-1.3.3/bk_installing_manually_book/content/rpm-chap-hue-5.html
(tomcatじゃないよ!!!)
画面はBootstrap, Knockout.js, jQuery
GitHub
https://github.com/cloudera/hue
Hue
http://gethue.com/
Hadoopの標準GUI HUEの最新情報
http://www.slideshare.net/Cloudera_jp/hadoopgui-hue
Hue SDK
http://cloudera.github.io/hue/docs-3.6.0/sdk/sdk.html
Hue関連のニュース
http://gethue-jp.tumblr.com/
インストール、設定
Install Hue without Cloudera
http://stackoverflow.com/questions/20579357/install-hue-without-cloudera
config
http://docs.hortonworks.com/HDPDocuments/HDP1/HDP-1.3.3/bk_installing_manually_book/content/rpm-chap-hue-5.html
2014年10月24日金曜日
hdfsで利用できる列指向フォーマット
現状ではRCFile だけが現実的。
ORCはHive専用状態。
Parquet は利用できるデータ型が中途半端。
(たぶん、半年後にはparquetがくるぞー)
Parquet,ORCFile についてのまとめページ
しかも日本語。
http://ozalog.blogspot.jp/2013/03/rcfileparquetorcfile.html
2014年10月23日木曜日
rails のSECRET_KEY_BASE でエラー
railsをproductionモードで起動したらエラー
Internal Server Error
Missing `secret_key_base` for 'production' environment, set this value in `config/secrets.yml`
config/secrets.yml
productionモード以外は固定値が入っている。rails4.1から?
とりあえずこれで動く。
# export SECRET_KEY_BASE=`bundle exec rake secret`
# rails s -e production
実際の運用ではどう設定しようか?
登録:
投稿 (Atom)