如何運行KafkaWordCount

如何運行KafkaWordCount，針對這個問題，這篇文章詳細介紹了相對應的分析和解答，希望可以幫助更多想解決這個問題的小伙伴找到更簡單易行的方法。

創(chuàng)新互聯(lián)公司專業(yè)提供四川服務器托管服務，為用戶提供五星數(shù)據(jù)中心、電信、雙線接入解決方案，用戶可自行在線購買四川服務器托管服務，并享受7*24小時金牌售后服務。

概要
Spark應用開發(fā)實踐性非常強，很多時候可能都會將時間花費在環(huán)境的搭建和運行上，如果有一個比較好的指導將會大大的縮短應用開發(fā)流程。Spark Streaming中涉及到和許多第三方程序的整合，源碼中的例子如何真正跑起來，文檔不是很多也不詳細。

下面主要講述如何運行KafkaWordCount，這個需要涉及Kafka集群的搭建，還是說的越仔細越好。

搭建Kafka集群
步驟1：下載kafka 0.8.1及解壓

wget https://www.apache.org/dyn/closer.cgi?path=/kafka/0.8.1.1/kafka_2.10-0.8.1.1.tgz
tar zvxf kafka_2.10-0.8.1.1.tgz
cd kafka_2.10-0.8.1.1
步驟2：啟動zookeeper

bin/zookeeper-server-start.sh config/zookeeper.properties
步驟3：修改配置文件config/server.properties，添加如下內容

host.name=localhost

# Hostname the broker will advertise to producers and consumers. If not set, it uses the
# value for "host.name" if configured. Otherwise, it will use the value returned from
# java.net.InetAddress.getCanonicalHostName().
advertised.host.name=localhost
步驟4：啟動Kafka server

bin/kafka-server-start.sh config/server.properties
步驟5：創(chuàng)建topic

bin/kafka-topics.sh --create --zookeeper localhost:2181 --replication-factor 1 --partitions 1 --topic test
檢驗topic創(chuàng)建是否成功

bin/kafka-topics.sh --list --zookeeper localhost:2181
如果正常返回test

步驟6：打開producer，發(fā)送消息

bin/kafka-console-producer.sh --broker-list localhost:9092 --topic test
##啟動成功后，輸入以下內容測試
This is a message
This is another message
步驟7：打開consumer，接收消息

bin/kafka-console-consumer.sh --zookeeper localhost:2181 --topic test --from-beginning
###啟動成功后，如果一切正常將會顯示producer端輸入的內容
This is a message
This is another message
運行KafkaWordCount
KafkaWordCount源文件位置examples/src/main/scala/org/apache/spark/examples/streaming/KafkaWordCount.scala

盡管里面有使用說明，見下文，但如果不是事先對Kafka有一定的了解的話，決然不知道這些參數(shù)是什么意思，也不知道該如何填寫。

/**
* Consumes messages from one or more topics in Kafka and does wordcount.
* Usage: KafkaWordCount
* is a list of one or more zookeeper servers that make quorum
* is the name of kafka consumer group
* is a list of one or more kafka topics to consume from
* is the number of threads the kafka consumer should use
*
* Example:
* `$ bin/run-example \
* org.apache.spark.examples.streaming.KafkaWordCount zoo01,zoo02,zoo03 \
* my-consumer-group topic1,topic2 1`
*/
object KafkaWordCount {
def main(args: Array[String]) {
if (args.length < 4) {
System.err.println("Usage: KafkaWordCount ")
System.exit(1)
}

StreamingExamples.setStreamingLogLevels()

val Array(zkQuorum, group, topics, numThreads) = args
val sparkConf = new SparkConf().setAppName("KafkaWordCount")
val ssc = new StreamingContext(sparkConf, Seconds(2))
ssc.checkpoint("checkpoint")

val topicpMap = topics.split(",").map((_,numThreads.toInt)).toMap
val lines = KafkaUtils.createStream(ssc, zkQuorum, group, topicpMap).map(_._2)
val words = lines.flatMap(_.split(" "))
val wordCounts = words.map(x => (x, 1L))
.reduceByKeyAndWindow(_ + _, _ - _, Minutes(10), Seconds(2), 2)
wordCounts.print()

ssc.start()
ssc.awaitTermination()
}
}

來看一看該如何運行KafkaWordCount

步驟1：停止運行剛才的kafka-console-producer和kafka-console-consumer

步驟2：運行KafkaWordCountProducer

bin/run-example org.apache.spark.examples.streaming.KafkaWordCountProducer localhost:9092 test 3 5
解釋一下參數(shù)的意思，localhost:9092表示producer的地址和端口, test表示topic，3表示每秒發(fā)多少條消息，5表示每條消息中有幾個單詞

步驟3：運行KafkaWordCount

bin/run-example org.apache.spark.examples.streaming.KafkaWordCount localhost:2181 test-consumer-group test 1
解釋一下參數(shù)， localhost:2181表示zookeeper的監(jiān)聽地址，test-consumer-group表示consumer-group的名稱，必須和$KAFKA_HOME/config/consumer.properties中的group.id的配置內容一致，test表示topic，1表示線程數(shù)。

關于如何運行KafkaWordCount問題的解答就分享到這里了，希望以上內容可以對大家有一定的幫助，如果你還有很多疑惑沒有解開，可以關注創(chuàng)新互聯(lián)行業(yè)資訊頻道了解更多相關知識。

分享文章：如何運行KafkaWordCount
文章分享：http://www.chinadenli.net/article6/giciig.html

成都網站建設公司_創(chuàng)新互聯(lián)，為您提供App開發(fā)、品牌網站建設、企業(yè)建站、動態(tài)網站、網站導航、外貿網站建設

聲明：本網站發(fā)布的內容（圖片、視頻和文字）以用戶投稿、用戶轉載內容為主，如果涉及侵權請盡快告知，我們將會在第一時間刪除。文章觀點不代表本網站立場，如需處理請聯(lián)系客服。電話：028-86922220；郵箱：631063699@qq.com。內容未經允許不得轉載，或轉載時需注明來源：創(chuàng)新互聯(lián)

猜你還喜歡下面的內容

欧美一区二区三区老妇人-欧美做爰猛烈大尺度电-99久久夜色精品国产亚洲a-亚洲福利视频一区二区

如何運行KafkaWordCount