Import files from Backblaze in Go: efficient techniques
This example imports a prefix from an existing Backblaze B2 bucket into a new private local directory. It uses the community Blazer client pinned to v0.5.3, with context-aware requests and a paginated iterator. It deliberately supports objects with a whole-file SHA-1 only.
Understanding the Backblaze B2 API
Blazer implements version 1 of the B2 Native API. This is an older third-party client, not an official Backblaze Go SDK. Backblaze continues to support older API versions, but multi-bucket application keys require version 4 authorization and cannot be used with this client. Use an existing version-1-compatible key for this maintenance example; choose a current client when your integration requires newer key types.
Backblaze’s API version and key compatibility policy
Setting up your Go environment for Backblaze file importing
Use Go 1.22 or newer and a fresh project directory. The B2 package in this version uses the standard library without additional runtime dependencies.
mkdir backblaze-import
cd backblaze-import
go mod init backblaze-import
go mod edit -go=1.22
GOTOOLCHAIN=local go get github.com/kurin/blazer/b2@v0.5.3
Using the Go SDK to connect to Backblaze B2
Set B2_KEY_ID to your application key ID and B2_APPLICATION_KEY to its secret using your secret
manager or shell environment. Give the key access to the existing bucket and the listBuckets,
listFiles and readFiles capabilities. The first credential argument is the key ID; it is not the
bucket ID. The program never creates a bucket.
Implementing basic file import functionality in Go
Save this complete program as main.go. Each download is checked against the size and SHA-1 from its
metadata, synced, closed and renamed within the same private directory. SHA-1 here checks transfer
integrity; it is not a signature proving who created the file.
package main
import (
"context"
"crypto/sha1"
"encoding/hex"
"errors"
"fmt"
"io"
"os"
"os/signal"
"path/filepath"
"strings"
"time"
"github.com/kurin/blazer/b2"
)
func importObject(ctx context.Context, object *b2.Object, directory string, number int) error {
attrs, err := object.Attrs(ctx)
if err != nil {
return err
}
digest, err := hex.DecodeString(attrs.SHA1)
if err != nil || len(digest) != sha1.Size || attrs.Size < 0 {
return errors.New("object needs a whole-file SHA-1; large files without one are unsupported")
}
reader := object.NewReader(ctx)
reader.ConcurrentDownloads = 1
reader.ChunkSize = 1 << 20
defer reader.Close()
tmp, err := os.CreateTemp(directory, ".import-")
if err != nil {
return err
}
defer os.Remove(tmp.Name())
defer tmp.Close()
hash := sha1.New()
n, err := io.Copy(io.MultiWriter(tmp, hash), io.LimitReader(reader, attrs.Size+1))
if err != nil {
return err
}
if n != attrs.Size || hex.EncodeToString(hash.Sum(nil)) != strings.ToLower(attrs.SHA1) {
return errors.New("object changed or download integrity check failed")
}
if err := reader.Close(); err != nil {
return err
}
if err := tmp.Sync(); err != nil {
return err
}
if err := tmp.Close(); err != nil {
return err
}
if err := ctx.Err(); err != nil {
return err
}
// Remote names never become local paths; even ../ and absolute keys stay contained.
name := fmt.Sprintf("%06d.bin", number)
if err := os.Rename(tmp.Name(), filepath.Join(directory, name)); err != nil {
return err
}
fmt.Printf("%s\t%q\t%d bytes\n", name, object.Name(), n)
return nil
}
func importPrefix(ctx context.Context, client *b2.Client, bucketName, prefix, parent string) (string, error) {
bucket, err := client.Bucket(ctx, bucketName)
if err != nil {
return "", err
}
if bucket == nil {
return "", errors.New("bucket lookup returned no bucket")
}
directory, err := os.MkdirTemp(parent, "b2-import-")
if err != nil {
return "", err
}
// Keep successful files after a later error. Never erase an earlier import.
iterator := bucket.List(ctx, b2.ListPrefix(prefix), b2.ListPageSize(100))
number := 0
for iterator.Next() {
if err := ctx.Err(); err != nil {
return directory, err
}
number++
if err := importObject(ctx, iterator.Object(), directory, number); err != nil {
return directory, fmt.Errorf("object %d: %w", number, err)
}
}
if err := iterator.Err(); err != nil {
return directory, err
}
return directory, ctx.Err()
}
func run() error {
if len(os.Args) != 4 {
return errors.New("usage: backblaze-import BUCKET PREFIX EXISTING_LOCAL_PARENT")
}
keyID, key := os.Getenv("B2_KEY_ID"), os.Getenv("B2_APPLICATION_KEY")
if keyID == "" || key == "" {
return errors.New("set B2_KEY_ID and B2_APPLICATION_KEY")
}
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt)
defer stop()
ctx, cancel := context.WithTimeout(ctx, 30*time.Minute)
defer cancel()
client, err := b2.NewClient(ctx, keyID, key)
if err != nil {
return err
}
directory, err := importPrefix(ctx, client, os.Args[1], os.Args[2], os.Args[3])
if directory != "" {
fmt.Fprintln(os.Stderr, "Import directory:", directory)
}
return err
}
func main() {
if err := run(); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
}
GOTOOLCHAIN=local go build -o backblaze-import .
./backblaze-import "${B2_BUCKET:?Set B2_BUCKET}" "${B2_PREFIX:-}" "${LOCAL_PARENT:?Set an existing local directory}"
Handling directories and batch imports: best practices
B2 object names are keys, not trusted filesystem paths. This example writes numbered .bin files and
prints a mapping from each local filename to its quoted remote key. An object named ../report or
/absolute/path therefore cannot escape the fresh import directory. The iterator fetches subsequent
pages automatically; iterator.Err() must be checked after iteration.
Error handling and debugging: ensuring smooth file transfers
Blazer handles its own HTTP retry classification and token refresh. Do not layer an unconditional retry loop around it: permission failures and missing buckets must be returned. All SDK calls share a cancellable context with a thirty-minute overall deadline, including the SDK’s backoff waits. A failed object’s temporary file is removed; earlier successful objects remain available.
Progress tracking for large file transfers
The program prints one completion record per verified file. It does not claim that bytes received so far are a completed import. For objects uploaded through B2’s large-file interface, a whole-file SHA-1 may be absent; this example fails explicitly instead of reporting unchecked data as verified.
Common pitfalls and troubleshooting tips
Keep the source prefix stable during the import. A listing is not an atomic snapshot of a changing bucket; size and checksum checks catch changed content during each download but do not freeze the listing itself.
Memory management
The SDK reader uses one concurrent download and 1 MiB chunks. Files stream to disk rather than accumulating in a byte slice. Only one object is processed at a time, and each reader is closed on both success and failure.
Rate limiting
Sequential object processing bounds request concurrency. The pinned client honors retry backoff for retryable responses; the overall context deadline bounds the total wait. If the account is throttled persistently, reduce workload or schedule another run after resolving the limit.
Handling network interruptions
Ctrl-C cancels the actual SDK requests and retry waits. Cancellation is not implemented by abandoning a goroutine while a download continues writing. A later run starts a new import directory rather than overwriting partial results from an earlier batch.
Concurrent downloads with rate limiting
Keep this batch loop sequential unless measurements justify parallel object imports. The SDK already
exposes bounded chunk concurrency through Reader.ConcurrentDownloads. Increasing it also increases
memory consumption and request pressure; it does not implement an account-wide request-rate limit.
Bucket lifecycle management best practices
Resolve an existing bucket before creating local output. Missing buckets and permission failures
return errors, and a nil bucket is rejected defensively. Bucket creation and lifecycle changes
belong in a separate provisioning step, so a typo during an import cannot create an unintended
resource.
Conclusion and further resources
This importer combines actual SDK cancellation and pagination with bounded streaming and verified local publication. It retains successful files after a later failure and reports the private import directory for inspection.
