Skip to content

Asset Upload

An asset is a file stored against one memory: the bytes live in object storage, and Hadron keeps a metadata row for them. Only a clean asset can be downloaded, and assets are scanned to get there — except on a self-hosted server running without a scanner and set to mark uploads clean (posture).

You can upload a file:

  • In the portal, by attaching it in an App's agent chat — see Upload a file in an agent chat.
  • With the CLI: hadron asset upload <file> -m <memory> and its siblings (list, get, url, rm, restore, link) — see hadron CLI → Assets.
  • Over GraphQL, with the memory-addressed flow below.

Not available today: agent-initiated uploads

An agent can't ask a user for a file: no upload-request tool is registered, and the portal shows no upload card. A file attached in the portal's agent chat isn't passed to the agent with the chat message. The older agent-addressed operations still exist and are deprecated — see The agent-addressed operations.

How an upload works

Three steps, and the bytes never pass through Hadron's API server:

  1. Begin — beginAssetUploadV2(input) with the target memoryId (an id or a fully-qualified memory URN), filename, mimeType, sizeBytes, uploadIntent and an optional description. It checks, in order: the memory exists; you can write to it (PERMISSION_DENIED); it accepts uploads (MEMORY_NOT_ACCEPTING_UPLOADS); the size (SIZE_EXCEEDED) and the type (MIME_NOT_ALLOWED) are within the limits. It returns {uploadId, putUrl, putHeaders, expiresAt, storageKey, maxSizeBytes, allowedMimeType}. No asset row exists yet.
  2. Upload — PUT the bytes to putUrl, sending putHeaders. The URL is signed for exactly the declared type and size, so object storage refuses anything else.
  3. Complete — completeAssetUpload({uploadId}), by the same caller, within 5 minutes of begin; after that the ticket has expired (NOT_FOUND, "upload ticket expired or unknown"). Complete checks the stored object's size and type, and — for types with a recognisable signature (PDF, images, Office files) — that the content matches the declared type (INVALID_INPUT on a mismatch; the staged bytes are deleted). It needs the memory's session key if the memory is encrypted (MEMORY_ENCRYPTED_NO_KEY), runs the scan (MALWARE_BLOCKED if it flags the file), and only then writes the row. It returns the Asset.

uploadIntent is required: pass USER. AGENT_REQUEST also requires an uploadRequestId (INVALID_INPUT otherwise) and belongs to the agent-request flow, which isn't wired up.

Which memories accept uploads

Memory.acceptsUploads is on by default. Memories created as an agent's system memory have it off. The field is read-only today: no mutation, portal setting or CLI command changes it.

The agent-addressed operations (deprecated)

Before the memory-addressed flow, uploads were addressed by agent: beginAssetUpload(agentUrn) put the file in the caller's per-agent memory, and agentAssets(agentUrn, …) listed it. Both still work and are deprecated — use beginAssetUploadV2 and memoryAssets.

They are opt-in per agent: the key asset:upload in the agentTools allowlist of the App that shares the agent's URN. Without it they return TOOL_NOT_ENABLED. The portal no longer has a control for this; the mutation setAgentAssetUploadEnabled remains, for org admins (PERMISSION_DENIED otherwise). The memory-addressed flow doesn't consult this key.

What the platform stores

The Asset type:

Field What it is
id The asset's id.
memoryId The memory it belongs to.
filename The original filename, verbatim.
mimeType, sizeBytes As declared at begin, and checked against the stored object at complete.
scanStatus PENDING, CLEAN or BLOCKED — lifecycle.
scanSignature The scanner's signature name on a BLOCKED asset; otherwise usually null. Read scanStatus, never infer it from this.
description Optional.
uploadedAt, uploadedBy When the row was written, and by whom (null for a file stored by a run).
deletedAt Set while soft-deleted.
urn The asset's URN, below.
publicUrl The public hotlink, or null.

The asset's URN is hrn:asset:<root>:<memory…>:assets:<id>. The memory part can be more than one segment, so read the id after the last assets segment rather than at a fixed position. Asset.urn can also come back in two degraded shapes:

  • hrn:asset:unknown:assets:<id>, when the holding memory couldn't be loaded; it carries no memory identity.
  • <memory-urn>:assets:<id>, the older form with no hrn:asset: prefix, when the memory's stored URN can't be rendered in the current grammar.

All three are accepted on input.

What the platform does not store in the row

  • The storage key. Bytes live under a key derived from the memory's URN and the asset id (…/<assetId>.<ext>), never from the filename, so no user-supplied, path-shaped text reaches storage. The key isn't exposed on the Asset type.
  • The bytes. Only object storage holds them; the row is metadata.

Listing

memoryAssets(memoryId, mimeType, skip = 0, count = 20, includeDeleted = false, mine = false) lists one memory's assets, newest first, up to 100 per page, as {assets, total, hasMore}. It needs read access to the memory (PERMISSION_DENIED). mine: true narrows to your own uploads. assets(filter, orgId, limit, offset) searches assets across memories.

Downloading

assetDownloadUrl(assetId, ttlSeconds = 300) returns a short-lived presigned URL, with the file's name, type and size. Default lifetime 5 minutes, maximum 1 hour. The URL isn't stored anywhere; ask for a fresh one each time. It needs read access to the asset's memory; an asset you can't read, or one that is deleted, is NOT_FOUND.

Downloads are gated on scanStatus = CLEAN:

  • PENDING returns SCAN_PENDING — try again in a few seconds. On a deployment with no scanner, or with the retry sweep off, retrying will not help: nothing advances the row.
  • BLOCKED returns SCAN_BLOCKED — the asset failed virus scanning and cannot be downloaded.
  • CLEAN returns the URL — or MEMORY_ENCRYPTED_NO_KEY if the memory is encrypted and has no session key (see Encrypted memories).

A server can also serve assets at an unauthenticated GET /assets/<id>: anyone holding the link gets a redirect to a 5-minute download. It is off by default; the operator turns it on with ASSET_PUBLIC_HOTLINK_ENABLED=true. Even then it serves only a clean, undeleted asset in a memory that isn't encrypted, and Asset.publicUrl is filled in only when the server also knows its public address (BASE_URL).

Every rejection — hotlinks off, unknown id, soft-deleted, not clean, encrypted — answers an identical bare 404. So a refusal never discloses which of those it was; it is not a claim that the endpoint hides the existence of assets it will happily serve.

The scan verdict lifecycle

scanStatus is a three-state contract (cor:dmo:060:12), and each state promises something specific to a caller.

State What it promises
PENDING Downloadable later, not an error. Fail-closed, and retried on backoff for as long as it takes. Never treat it as terminal.
CLEAN Scanned and settled. The only state that downloads.
BLOCKED Terminal. No later verdict overwrites it.

An infected upload fails at upload time, not on first download: the call returns the typed MALWARE_BLOCKED error, so the caller learns immediately.

BLOCKED is a tombstone, not a quarantine. The row is kept as the audit record and carries the engine's signature name in Asset.scanSignature, which is readable through the API — but the object bytes are deleted.

Deleted, not never-written: the scan runs against the bytes at their final location, so infected bytes exist in the store between the upload landing and the verdict, and longer if the delete fails — it is retried until the store confirms. What the tombstone guarantees is that they are not retained and never become downloadable, not that they were never there.

That retry is the same background sweep that advances PENDING rows, so DISABLE_ASSET_SCAN_SWEEP suspends it too — on a deployment with the sweep off, a delete that fails is not retried and the infected bytes stay until an operator removes them. It is the one setting on this page with a malware-retention consequence.

Scanner unavailability never fails an upload — only a real verdict does. A verdict is accepted only from a scan capability tool declaring the versioned scan@1 contract; an answer without it (a misrouted URL, an incompatible implementation) counts as unavailable, never as clean. So a misconfigured scanner leaves assets pending rather than waving them through.

Scans always run against the bytes at their final storage location, never one the uploader can still write to.

Indefinite retry is a platform promise, not a deployment guarantee

The contract is that nothing gives up on a PENDING row. Two operator choices suspend it, and both are self-hosting decisions rather than platform behaviour: DISABLE_ASSET_SCAN_SWEEP turns the retry sweep off, and a deployment with no scanner configured has nothing to retry with. In either case rows stay pending — still fail-closed, still undownloadable, but no longer on their way anywhere.

The same sweep retries BLOCKED byte purges, so disabling it also strands infected bytes whose first delete failed.

On the managed deployment the promise holds as written.

Limits

These are what the server enforces, and they are platform-wide: the upload operation accepts no per-call or per-agent limits, so nothing raises or lowers them. (A per-App override existed and was removed in hadron-server#900 — every call site passed null, so it had never resolved to anything but the platform default.)

  • Maximum size: 25 MB.
  • Allowed MIME types: application/pdf; image/png, image/jpeg, image/webp; text/plain, text/markdown, text/csv, application/json; and the OOXML trio — .docx, .xlsx, .pptx. GIF is not on the list.
  • Begin to complete: 5 minutes — the upload URL and the ticket expire together.
  • Download URL lifetime: 5 minutes by default, 1 hour maximum.

Soft-delete and restore

softDeleteAsset(assetId) marks an asset deleted: deletedAt is set and it drops out of listings. restoreAsset(assetId) undoes it within 24 hours. Only the uploader or an org admin may do either (PERMISSION_DENIED). Neither works yet for an asset in a user-owned memory (NOT_FOUND).

An hourly janitor purges soft-deleted assets older than 24 hours: the stored object first, then the row. Restoring after that — or when the bytes are already gone — returns ASSET_BYTES_GONE, rather than re-creating an empty row. Restoring an asset that isn't deleted returns ASSET_NOT_DELETED.

Encrypted memories

In an encrypted memory, the asset's storage key is encrypted at rest with the memory's key (AES-256-GCM), and both uploading and downloading need the memory's session key. Without it, the operation returns MEMORY_ENCRYPTED_NO_KEY; the portal says the memory is encrypted and has no active session key, and asks you to unlock it and try again. Assets in encrypted memories are never served by the public hotlink.

Linking a file into the graph

createAssetReferenceNode(input) creates a node that references an asset, so a file can take part in the knowledge graph; the node may live in a different memory from the asset. hadron asset link <asset-ref> --node <new-node-urn> does the same from the CLI. It's a separate, deliberate step — nothing offers it at upload time.

Running the scanner (self-hosted)

Hadron's managed deployment scans uploads for you; this section is for operators running their own. The engine is a provider detail — the platform talks to a scan capability tool over the scan@1 contract, and ClamAV is what sits behind it today.

Core reads:

Variable Purpose
SCAN_TOOL_URL The scan tool's base URL, e.g. http://hadrontool-scan:8080. Unset means no scanner — see the posture note below.
SCAN_TOOL_TOKEN Shared bearer token, sent when set.
DISABLE_ASSET_SCAN_SWEEP Turns off the background sweep that retries PENDING scans and BLOCKED byte purges.
HADRON_ASSET_SCAN_AUTO_CLEAN The scanner-less hatch — see below.
ASSET_PUBLIC_HOTLINK_ENABLED true turns on the public hotlink. Off by default.
DISABLE_ASSET_JANITOR true turns off the hourly sweep that purges soft-deleted assets after 24 hours.
ASSET_ORPHAN_REAP_ENABLED true lets the daily reaper delete abandoned staging objects older than 24 hours; otherwise it only reports them.

The hadrontool-scan sidecar wraps clamd and reads SCAN_TOOL_TOKEN, CLAMD_HOST, CLAMD_PORT, SCAN_MAX_BYTES, and SCAN_ALLOW_PRIVATE_NETWORKS.

Deploying ClamAV alongside it (clamav/clamav:1.4), three settings carry most of the operational weight:

  • StreamMaxLength ≥ the platform's 25 MB cap. Below it, the largest uploads the platform accepts are the ones the scanner refuses.
  • ConcurrentDatabaseReload=no holds memory at the ~1.5 GB signature floor. With it on, a reload transiently doubles that.
  • SCAN_ALLOW_PRIVATE_NETWORKS=true when the object store is on a private address — a local MinIO, for instance. Otherwise the sidecar refuses to fetch the bytes it was asked to scan.

Running without a scanner is a posture, not a default

With no SCAN_TOOL_URL, uploads stay PENDING — and therefore undownloadable, because the clean-only gate does not relax. That is fail-closed and deliberate.

The escape hatch is HADRON_ASSET_SCAN_AUTO_CLEAN=true, which marks uploads born-clean. Understand what you are choosing: uploads are then served unscanned.

The hatch is ignored whenever a scanner is configured, so the scanner always wins and turning one on later is order-safe — you cannot leave the hatch enabled and silently keep bypassing a scanner you have since deployed.

One server instance

An upload's ticket — the link between beginAssetUploadV2 and completeAssetUpload — is held in the server process's memory. Both calls must reach the same instance, within the 5-minute window. Running several instances behind a load balancer without affinity breaks uploads until the ticket is moved out of process.

What's deferred

  • Agent-initiated upload requests — see the note at the top.
  • Content extraction. OCR for PDFs and images, with extracted text saved on the asset for retrieval.
  • Cross-memory sharing. An asset belongs to one memory. Searching across memories exists (assets); sharing one asset with other users or agents needs explicit policy.
  • Customer-managed encryption keys for the bytes themselves; today's bytes use the object store's server-side encryption.
  • Versioning. Replacing a file doesn't keep the prior version.