第7课：Spark Streaming源码解读之JobScheduler内幕实现和深度思考

时间 2019-11-18

标签 spark streaming 源码解读 jobscheduler 内幕实现深度思考栏目 Spark 繁體版

原文原文链接

本期内容：app

1，JobScheduler内幕实现ide

2，JobScheduler深度思考函数

DStream的foreachRDD方法，实例化ForEachDStream对象，并将用户定义的函数foreachFunc传入到该对象中。foreachRDD方法是输出操做，foreachFunc方法会做用到这个DStream中的每一个RDD。this

/**
* Apply a function to each RDD in this DStream. This is an output operator, so
* 'this' DStream will be registered as an output stream and therefore materialized.
* @param foreachFunc foreachRDD function
* @param displayInnerRDDOps Whether the detailed callsites and scopes of the RDDs generated
*                           in the `foreachFunc` to be displayed in the UI. If `false`, then
*                           only the scopes and callsites of `foreachRDD` will override those
*                           of the RDDs on the display.
*/
private def foreachRDD(
    foreachFunc: (RDD[T], Time) => Unit,
    displayInnerRDDOps: Boolean): Unit = {
  new ForEachDStream(this,
    context.sparkContext.clean(foreachFunc, false), displayInnerRDDOps).register()
}spa

ForEachDStream对象中重写了generateJob方法，调用父DStream的getOrCompute方法来生成RDD并封装Job，传入对该RDD的操做函数foreachFunc和time。dependencies方法定义为父DStream的集合。.net

/**
* An internal DStream used to represent output operations like DStream.foreachRDD.
* @param parent        Parent DStream
* @param foreachFunc   Function to apply on each RDD generated by the parent DStream
* @param displayInnerRDDOps Whether the detailed callsites and scopes of the RDDs generated
*                           by `foreachFunc` will be displayed in the UI; only the scope and
*                           callsite of `DStream.foreachRDD` will be displayed.
*/
private[streaming]
class ForEachDStream[T: ClassTag] (
    parent: DStream[T],
    foreachFunc: (RDD[T], Time) => Unit,
    displayInnerRDDOps: Boolean
  ) extends DStream[Unit](parent.ssc) {

  override def dependencies: List[DStream[_]] = List(parent)

  override def slideDuration: Duration = parent.slideDuration

  override def compute(validTime: Time): Option[RDD[Unit]] = None

  override def generateJob(time: Time): Option[Job] = {
    parent.getOrCompute(time) match {
      case Some(rdd) =>
        val jobFunc = () => createRDDWithLocalProperties(time, displayInnerRDDOps) {
          foreachFunc(rdd, time)
        }
        Some(new Job(time, jobFunc))
      case None => None
    }
  }
}线程

DStreamGraph的generateJobs方法中会调用outputStream的generateJob方法，就是调用ForEachDStream的generateJob方法。scala

def generateJobs(time: Time): Seq[Job] = {
  logDebug("Generating jobs for time " + time)
  val jobs = this.synchronized {
    outputStreams.flatMap { outputStream =>
      val jobOption = outputStream.generateJob(time)
      jobOption.foreach(_.setCallSite(outputStream.creationSite))
      jobOption
    }
  }
  logDebug("Generated " + jobs.length + " jobs for time " + time)
  jobs
}对象

DStream的generateJob定义以下，其子类中只有ForEachDStream重写了generateJob方法。ci

/**
* Generate a SparkStreaming job for the given time. This is an internal method that
* should not be called directly. This default implementation creates a job
* that materializes the corresponding RDD. Subclasses of DStream may override this
* to generate their own jobs.
*/
private[streaming] def generateJob(time: Time): Option[Job] = {
  getOrCompute(time) match {
    case Some(rdd) => {
      val jobFunc = () => {
        val emptyFunc = { (iterator: Iterator[T]) => {} }
        context.sparkContext.runJob(rdd, emptyFunc)
      }
      Some(new Job(time, jobFunc))
    }
    case None => None
  }
}

DStream的print方法内部仍是调用foreachRDD来实现，传入了内部方法foreachFunc，来取出num+1个数后打印输出。

/**
* Print the first num elements of each RDD generated in this DStream. This is an output
* operator, so this DStream will be registered as an output stream and there materialized.
*/
def print(num: Int): Unit = ssc.withScope {
  def foreachFunc: (RDD[T], Time) => Unit = {
    (rdd: RDD[T], time: Time) => {
      val firstNum = rdd.take(num + 1)
      // scalastyle:off println
      println("-------------------------------------------")
      println("Time: " + time)
      println("-------------------------------------------")
      firstNum.take(num).foreach(println)
      if (firstNum.length > num) println("...")
      println()
      // scalastyle:on println
    }
  }
  foreachRDD(context.sparkContext.clean(foreachFunc), displayInnerRDDOps = false)
}

总结：JobScheduler是SparkStreaming 全部Job调度的中心，内部有两个重要的成员：

JobGenerator负责Job的生成，ReceiverTracker负责记录输入的数据源信息。

JobScheduler的启动会致使ReceiverTracker和JobGenerator的启动。ReceiverTracker的启动致使运行在Executor端的Receiver启动而且接收数据，ReceiverTracker会记录Receiver接收到的数据meta信息。JobGenerator的启动致使每隔BatchDuration，就调用DStreamGraph生成RDD Graph，并生成Job。JobScheduler中的线程池来提交封装的JobSet对象(时间值，Job，数据源的meta)。Job中封装了业务逻辑，致使最后一个RDD的action被触发，被DAGScheduler真正调度在Spark集群上执行该Job。