> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tuneplane.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Job Events

> Answers "what happened before the job started": backend events -- image pulls, failed scheduling, container exits.

The split with /logs: a log is written by the training process, so it is empty until
that process starts -- and the hardest window to diagnose is exactly that one. An image
pulling (tens of GB, tens of minutes), a Pod that will not schedule, a container that
exits as soon as it starts. On kuberay the real reason for "Pending forever" is **only
in the K8s events**, RayJob.status says nothing at all; on an agent Fleet, "its GPUs
were taken while the image pulled" is known only to the node daemon.

A job not yet handed to a backend (QUEUED or PAUSED, no job_ref) has no backend events
to show -- "why has it not started" is /scheduling's question, and these two endpoints
divide the work that way.



## OpenAPI

````yaml /api-reference/openapi.json get /api/jobs/{run_id}/events
openapi: 3.1.0
info:
  title: TunePlane Console
  description: >-
    The TunePlane control plane. Everything the `tuneplane` CLI and the web
    console do goes through this API, and so can your own tooling.


    Authenticate with a bearer token from `POST /api/auth/login` or a CLI device
    flow; see the Authentication page for how to get one and how long it lasts.
  version: 0.3.42
servers:
  - url: https://{host}
    description: Your TunePlane deployment
    variables:
      host:
        default: tuneplane.your-company.com
        description: >-
          The domain your administrator gave you, without a scheme or trailing
          slash.
security: []
tags:
  - name: auth
    description: >-
      Log in, exchange a CLI device code, and inspect the current identity.
      Everything else on this API needs a bearer token from here.
  - name: profile
    description: >-
      The signed-in user's own account: quota, tokens, preferences, and
      notification settings.
  - name: projects
    description: >-
      Projects group runs the way `tuneplane.yaml` names them. A run belongs to
      exactly one.
  - name: experiments
    description: >-
      Read the experiment definitions the console found in the configured
      repository.
  - name: submit
    description: >-
      Admit a JobSpec. This is what `tp submit` calls: the catalog handshake,
      quota check, and preflight all happen here, and a rejection names the gate
      that refused it.
  - name: jobs
    description: >-
      Everything about a job after it is admitted: status, logs, metrics,
      samples, artifacts, and the pause/resume/stop controls.
  - name: runs
    description: >-
      Finished work, addressed by run id. A run outlives the job that produced
      it.
  - name: ingest
    description: >-
      The endpoints training code reports to. `tuneplane.report` speaks this;
      you only call it directly when writing an adapter for a framework the
      catalog does not cover.
  - name: datasets
    description: >-
      Versioned dataset upload, listing, and metadata. Protected datasets expose
      identity and schema here but never their records.
  - name: volumes
    description: Governed directories of files a job may mount read-only.
  - name: environments
    description: >-
      Agent RL environments: their manifests, versions, and upload URLs. A
      taskset is never returned.
  - name: benchmarks
    description: >-
      The benchmark catalog, the score matrix across runs, and externally scored
      evaluations.
  - name: rubrics
    description: >-
      Written scoring standards, their revisions, and which runs cited which
      version.
  - name: judge
    description: >-
      The LLM-judge endpoint a training job calls to score a rollout.
      OpenAI-compatible.
  - name: models
    description: >-
      The model registry: register a version, promote it, archive it, read its
      card.
  - name: model-deployments
    description: >-
      Managed model versions serving application traffic: revisions, promotion,
      rollback, suspension, and deployment tokens.
  - name: inference
    description: >-
      OpenAI-compatible inference against a promoted deployment revision. This
      is the endpoint applications call.
  - name: playground
    description: >-
      Short-lived serving sessions for human evaluation. Distinct from a
      deployment: a session expires, a deployment does not.
  - name: reflow
    description: >-
      The governed path from a deployment's production traffic back to the
      training data of its next version.
  - name: annotate
    description: 'Preference annotation: pull a batch, push judgements, read progress.'
  - name: plugins
    description: Installed plugins and the extension shelf the console renders.
  - name: diagnosis
    description: >-
      Automated analysis of a finished or failed run, and the accumulated
      project memory it draws on.
  - name: approvals
    description: 'Approval requests: an escalation path, one level deep, with a record.'
  - name: billing
    description: >-
      What the GPU-hours cost. One price on top of the hours the usage page
      already shows.
  - name: teams
    description: 'Teams: the unit capacity is budgeted to. A department, not a tenant.'
  - name: agent
    description: >-
      Submit plans: a proposed submission a human approves or rejects before it
      becomes a job.
  - name: share
    description: >-
      Public, revocable read-only links to a job or a comparison. The
      `/api/share/{token}` routes need no bearer token, which is the point.
  - name: notifications
    description: The signed-in user's notification feed.
  - name: search
    description: Cross-surface search over jobs, runs, datasets, and models.
  - name: sandbox
    description: >-
      Execute model-generated code in a throwaway container with no GPU and no
      network.
  - name: uploads
    description: >-
      Resumable upload sessions used by dataset, environment, and plugin
      publishing.
  - name: integrations-hf
    description: Hugging Face account linking and repository push.
  - name: mcp
    description: Model Context Protocol access information and per-user tool settings.
  - name: mcp-oauth
    description: >-
      OAuth metadata, authorization, token exchange, and dynamic client
      registration for MCP clients.
  - name: cluster
    description: Live capacity and node state across the fleets.
  - name: fleets
    description: >-
      Registered execution backends and the machines in them. Reading is open to
      every user; creating a fleet, minting a join token and draining a node are
      admin-only. Joining is authorized by the join token alone.
  - name: admin
    description: >-
      User, role, quota, hardware, schedule, integration, and settings
      administration. Admin role required.
  - name: tasks
    description: >-
      Scheduled platform maintenance tasks: what they are, when they last ran,
      and running one now.
  - name: report
    description: The rendered daily report page.
  - name: health
    description: Liveness and version. Unauthenticated.
paths:
  /api/jobs/{run_id}/events:
    get:
      tags:
        - jobs
      summary: Job Events
      description: >-
        Answers "what happened before the job started": backend events -- image
        pulls, failed scheduling, container exits.


        The split with /logs: a log is written by the training process, so it is
        empty until

        that process starts -- and the hardest window to diagnose is exactly
        that one. An image

        pulling (tens of GB, tens of minutes), a Pod that will not schedule, a
        container that

        exits as soon as it starts. On kuberay the real reason for "Pending
        forever" is **only

        in the K8s events**, RayJob.status says nothing at all; on an agent
        Fleet, "its GPUs

        were taken while the image pulled" is known only to the node daemon.


        A job not yet handed to a backend (QUEUED or PAUSED, no job_ref) has no
        backend events

        to show -- "why has it not started" is /scheduling's question, and these
        two endpoints

        divide the work that way.
      operationId: job_events_api_jobs__run_id__events_get
      parameters:
        - name: run_id
          in: path
          required: true
          schema:
            type: string
            title: Run Id
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/JobEventsOut'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    JobEventsOut:
      properties:
        events:
          items:
            $ref: '#/components/schemas/JobEventOut'
          type: array
          title: Events
          default: []
        backend:
          type: string
          title: Backend
          default: ''
        backend_unavailable:
          type: boolean
          title: Backend Unavailable
          default: false
      type: object
      title: JobEventsOut
      description: >-
        A job's events.


        backend_unavailable=True means this round could not reach the backend at
        all -- a

        network blip, a node out of contact -- which is a different thing from
        there being

        no events. The console has to tell them apart, or a reader takes
        "nothing here" for

        an answer the platform actually obtained.
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    JobEventOut:
      properties:
        type:
          type: string
          title: Type
          default: Normal
        reason:
          type: string
          title: Reason
          default: ''
        message:
          type: string
          title: Message
          default: ''
        count:
          anyOf:
            - type: integer
            - type: 'null'
          title: Count
        last:
          anyOf:
            - type: string
            - type: 'null'
          title: Last
      type: object
      title: JobEventOut
      description: >-
        One job event from the execution backend.


        Before the process comes up the logs are empty and the status is a bare
        PENDING;

        the reason is only here -- the image is still pulling, the pod will not
        schedule,

        the container starts and exits.
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
        input:
          title: Input
        ctx:
          type: object
          title: Context
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````