Sincronización concurrente de archivos con Go y MinIO
Descarga un bucket de MinIO en un directorio local con cinco workers de Go. El ejemplo conserva las rutas anidadas, prepara cada archivo antes de reemplazar una copia existente y devuelve un error si una transferencia falla o presionas Ctrl+C. Es un descargador unidireccional: cada ejecución vuelve a descargar todos los objetos y las eliminaciones remotas no borran archivos locales.
Configura el servidor MinIO
Usa Bash en Linux con un sistema de archivos local que distinga mayúsculas de minúsculas,
Go 1.26.8 en tu
PATH, y curl y diff instalados.
Este tutorial se probó con Go 1.26.8 y MinIO Go SDK v7.0.95. Los puertos 9000 y 9001 deben estar libres.
Al 3 de octubre de 2026, el repositorio comunitario de MinIO está archivado y ya no recibe mantenimiento. Su distribución comunitaria solo ofrece código fuente. El servidor con versión fijada que se muestra a continuación es un entorno local aislado para aprender, no una recomendación para un nuevo despliegue en producción. Utiliza la instalación desde código fuente publicada por MinIO para esa versión, en lugar de una imagen de contenedor sin versión fijada.
Pega lo siguiente en tu shell desde un directorio donde quieras crear un directorio
minio-sync nuevo. El subshell mantiene intactos tu directorio actual y las
opciones del shell; si el destino ya existe, la configuración se detiene antes de la instalación.
Compilar el servidor puede tardar varios minutos:
(
set -e
mkdir minio-sync
cd minio-sync
GOBIN="$PWD/bin" GOWORK=off GOTOOLCHAIN=local \
go install github.com/minio/minio@RELEASE.2025-10-15T17-29-55Z
)
Inicia el servidor en esa terminal y déjalo en ejecución:
(
cd minio-sync &&
MINIO_ROOT_USER=minioadmin MINIO_ROOT_PASSWORD=minioadmin \
./bin/minio server ./data --address 127.0.0.1:9000 --console-address 127.0.0.1:9001
)
Ambos puntos de escucha se vinculan a la interfaz de loopback. Estas credenciales conocidas y
HTTP sin cifrar solo deben usarse en este entorno de pruebas privado. Presiona Ctrl+C en esta
terminal cuando termines; el servidor se detiene y su directorio data
queda disponible para otra ejecución.
Inicializa el proyecto
En una segunda terminal, comienza desde el mismo directorio padre. Confirma que el servidor esté
listo; luego inicializa el módulo e instala la versión fijada del SDK.
GOWORK=off evita que un espacio de trabajo de Go contenedor cambie la selección
de módulos del ejemplo:
(
set -e
curl -fsS http://127.0.0.1:9000/minio/health/ready
cd minio-sync
GOWORK=off GOTOOLCHAIN=local go mod init minio-sync
GOWORK=off GOTOOLCHAIN=local go get github.com/minio/minio-go/v7@v7.0.95
)
Llena el bucket
Crea minio-sync/_seed/ y guarda lo siguiente como _seed/main.go.
Escribe archivos de entrada conocidos en sample-input/ y los sube a
my-sync-bucket: 300 archivos de texto, un archivo de texto anidado, un archivo vacío
y un archivo binario. El directorio con guion bajo mantiene este ejecutable independiente fuera
del paquete principal.
// _seed/main.go
package main
import (
"context"
"fmt"
"log"
"os"
"path/filepath"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio-go/v7/pkg/credentials"
)
func main() {
if err := seed(); err != nil {
log.Fatal(err)
}
log.Println("Uploaded 303 sample objects")
}
func seed() error {
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
defer cancel()
client, err := minio.New("localhost:9000", &minio.Options{
Creds: credentials.NewStaticV4("minioadmin", "minioadmin", ""),
Secure: false,
})
if err != nil {
return err
}
const bucket = "my-sync-bucket"
exists, err := client.BucketExists(ctx, bucket)
if err != nil {
return err
}
if !exists {
if err := client.MakeBucket(ctx, bucket, minio.MakeBucketOptions{}); err != nil {
return err
}
}
files := map[string][]byte{
"notes/hello.txt": []byte("Hello from MinIO\n"),
"empty.bin": {},
"binary.bin": {0, 255, 128, 13, 10, 42, 0, 10},
}
for i := 0; i < 300; i++ {
files[fmt.Sprintf("logs/object-%03d.txt", i)] = []byte(fmt.Sprintf("object %d\n", i))
}
for key, body := range files {
filename := filepath.Join("sample-input", filepath.FromSlash(key))
if err := os.MkdirAll(filepath.Dir(filename), 0o700); err != nil {
return err
}
if err := os.WriteFile(filename, body, 0o600); err != nil {
return err
}
if _, err := client.FPutObject(ctx, bucket, key, filename, minio.PutObjectOptions{}); err != nil {
return fmt.Errorf("upload %q: %w", key, err)
}
}
return nil
}
Ejecuta el programa de carga inicial desde el directorio padre:
(
cd minio-sync &&
GOWORK=off GOTOOLCHAIN=local go run ./_seed
)
Informa Uploaded 303 sample objects. Volver a ejecutarlo reemplaza esas claves de ejemplo y los
archivos de entrada; úsalo solo con este servidor y directorio dedicados. Mantén sin cambios los
datos del servidor y sample-input/ mientras descargas y comparas los resultados.
Configura el cliente MinIO
Guarda lo siguiente como minio-sync/client.go. El servidor local usa HTTP, por lo que
useSSL es false.
La versión mínima de TLS del transporte solo se aplica cuando useSSL
es true y el endpoint sirve HTTPS.
BucketExists debe devolver tanto un valor verdadero como la ausencia de errores
antes de que comience la descarga:
// client.go
package main
import (
"context"
"crypto/tls"
"fmt"
"net/http"
"time"
"github.com/minio/minio-go/v7"
"github.com/minio/minio-go/v7/pkg/credentials"
)
// bucketName matches the bucket populated by _seed/main.go.
const bucketName = "my-sync-bucket"
func createMinioClient(ctx context.Context) (*minio.Client, error) {
endpoint := "localhost:9000"
accessKeyID := "minioadmin" // Use environment variables in production
secretAccessKey := "minioadmin" // Use environment variables in production
useSSL := false // true once your endpoint terminates TLS
// Only exercised when useSSL is true
transport := &http.Transport{
TLSClientConfig: &tls.Config{MinVersion: tls.VersionTLS12},
IdleConnTimeout: 90 * time.Second,
}
// Initialize minio client
opts := &minio.Options{
Creds: credentials.NewStaticV4(accessKeyID, secretAccessKey, ""),
Secure: useSSL,
Transport: transport,
}
client, err := minio.New(endpoint, opts)
if err != nil {
return nil, err
}
// BucketExists reports a bucket that is simply absent as (false, nil), so the boolean
// has to be checked too. Ignoring it turns a typo in the bucket name into a sync that
// quietly downloads nothing.
exists, err := client.BucketExists(ctx, bucketName)
if err != nil {
return nil, fmt.Errorf("failed to reach bucket %q: %w", bucketName, err)
}
if !exists {
return nil, fmt.Errorf("bucket %q does not exist", bucketName)
}
return client, nil
}
Implementa descargas concurrentes de archivos
Guarda lo siguiente como minio-sync/sync.go. Cinco workers leen tareas mientras una
goroutine independiente consume los resultados; esperar a que termine el listado para consumir
los resultados provocaría un interbloqueo con un bucket que supere la capacidad de los búferes de
los canales. Cada objeto tiene un tiempo de espera máximo de 10 minutos. Se cuentan los errores y
se incluye la clave del primer objeto que falla en el diagnóstico final.
La referencia de GetObject del SDK
explica que los errores suelen llegar al leer el flujo. Por eso, comprobar
io.Copy importa tanto como comprobar GetObject.
Una transferencia fallida elimina su archivo temporal y deja intacta la copia completada
anteriormente.
// sync.go
package main
import (
"context"
"fmt"
"io"
"os"
"path"
"path/filepath"
"strings"
"sync"
"time"
"github.com/minio/minio-go/v7"
)
type DownloadResult struct {
ObjectName string
Error error
}
func downloadFiles(ctx context.Context, client *minio.Client, bucketName string, outputDir string) error {
// 0700, because these are private copies. The process umask would otherwise usually
// leave them group- and world-readable.
if err := os.MkdirAll(outputDir, 0o700); err != nil {
return fmt.Errorf("failed to create output directory: %w", err)
}
// Resolve the directory once so the per-object checks below compare against a path
// that has no symlinks left in it.
root, err := filepath.EvalSymlinks(outputDir)
if err != nil {
return fmt.Errorf("failed to resolve output directory: %w", err)
}
// Cancelling here unblocks every worker if we bail out early
ctx, cancel := context.WithCancel(ctx)
defer cancel()
jobs := make(chan string, 100)
results := make(chan DownloadResult, 100)
var wg sync.WaitGroup
workerCount := 5
// Start workers
for i := 0; i < workerCount; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for objectName := range jobs {
err := downloadObject(ctx, client, bucketName, objectName, root)
select {
case results <- DownloadResult{ObjectName: objectName, Error: err}:
case <-ctx.Done():
return
}
}
}()
}
go func() {
wg.Wait()
close(results)
}()
// Collect results while the producer is still queueing. Draining only after
// close(jobs) deadlocks as soon as the buffers fill: workers block on
// `results <-`, stop reading `jobs`, and the producer blocks on `jobs <-`.
//
// Only a count and the first error are kept. A bucket can hold millions of objects,
// and appending every failure to a slice would grow without bound in exactly the run
// where things are going worst.
type failures struct {
count int
first error
}
collected := make(chan failures, 1)
go func() {
var seen failures
for result := range results {
if result.Error == nil {
continue
}
seen.count++
if seen.first == nil {
seen.first = fmt.Errorf("%s: %w", result.ObjectName, result.Error)
}
}
collected <- seen
}()
// List and queue objects. The deferred close lets the workers drain and exit
// on every path, including the early returns below.
listErr := func() error {
defer close(jobs)
opts := minio.ListObjectsOptions{Recursive: true}
for obj := range client.ListObjects(ctx, bucketName, opts) {
if obj.Err != nil {
// Cancel immediately so in-flight downloads stop now rather than running
// to completion behind a listing that has already failed.
cancel()
return fmt.Errorf("error listing objects: %w", obj.Err)
}
// Empty directory markers contain no file bytes to download.
if obj.Size == 0 && strings.HasSuffix(obj.Key, "/") {
continue
}
select {
case jobs <- obj.Key:
case <-ctx.Done():
return ctx.Err()
}
}
return nil
}()
// Always drain first, so every worker has finished before we report anything.
seen := <-collected
if listErr != nil {
return listErr
}
// A cancelled context ends the listing loop without an error of its own, and workers
// that were cut off never deliver a result. Without this check a cancelled sync would
// be indistinguishable from an empty bucket that synced cleanly.
if err := ctx.Err(); err != nil {
return fmt.Errorf("sync did not complete: %w", err)
}
if seen.count > 0 {
return fmt.Errorf("encountered %d download errors, first: %w", seen.count, seen.first)
}
return nil
}
// resolveOutputPath maps a server-supplied object key onto a path inside outputDir, which
// must already be symlink-free (see filepath.EvalSymlinks above).
//
// Keys are rejected rather than cleaned. `a/./b.txt` and `a/b/../b.txt` are different keys
// that both clean to `a/b.txt`, so cleaning them would let one object silently overwrite
// another, and lexical checks alone cannot see a symlink that is already on disk.
func resolveOutputPath(outputDir, objectName string) (string, error) {
if objectName == "" || objectName != path.Clean(objectName) || strings.Contains(objectName, `\`) {
return "", fmt.Errorf("refusing non-canonical object key: %q", objectName)
}
if !filepath.IsLocal(filepath.FromSlash(objectName)) {
return "", fmt.Errorf("refusing unsafe object key: %q", objectName)
}
current := outputDir
for _, element := range strings.Split(objectName, "/") {
current = filepath.Join(current, element)
info, err := os.Lstat(current)
if err != nil {
if os.IsNotExist(err) {
continue // Nothing here yet, so nothing can redirect the write
}
return "", fmt.Errorf("failed to inspect %q: %w", current, err)
}
if info.Mode()&os.ModeSymlink != 0 {
return "", fmt.Errorf("refusing object key through a symlink: %q", objectName)
}
}
return current, nil
}
func downloadObject(ctx context.Context, client *minio.Client, bucket, objectName, outputDir string) error {
outputPath, err := resolveOutputPath(outputDir, objectName)
if err != nil {
return err
}
// Create context with timeout
ctx, cancel := context.WithTimeout(ctx, 10*time.Minute)
defer cancel()
// Get object
obj, err := client.GetObject(ctx, bucket, objectName, minio.GetObjectOptions{})
if err != nil {
return fmt.Errorf("failed to get object: %w", err)
}
defer obj.Close()
if err := os.MkdirAll(filepath.Dir(outputPath), 0o700); err != nil {
return fmt.Errorf("failed to create directories: %w", err)
}
// Download to a sibling temp file so a failed sync never truncates the copy
// that a previous run completed, and never leaves a half-written file behind.
// os.CreateTemp already creates the file 0600.
temp, err := os.CreateTemp(filepath.Dir(outputPath), filepath.Base(outputPath)+".part-*")
if err != nil {
return fmt.Errorf("failed to create temporary file: %w", err)
}
tempPath := temp.Name()
defer func() {
temp.Close()
os.Remove(tempPath) // No-op once the rename below succeeded
}()
if _, err := io.Copy(temp, obj); err != nil {
return fmt.Errorf("failed to download file: %w", err)
}
if err := temp.Sync(); err != nil {
return fmt.Errorf("failed to flush file: %w", err)
}
if err := temp.Close(); err != nil {
return fmt.Errorf("failed to close file: %w", err)
}
// Rename is atomic within a filesystem, so readers see either the old file or
// the complete new one
if err := os.Rename(tempPath, outputPath); err != nil {
return fmt.Errorf("failed to publish file: %w", err)
}
return nil
}
Manejo de errores y resolución de problemas
Se omiten los marcadores de directorios vacíos, que tienen cero bytes y una clave terminada en
/. Un archivo vacío normal sí se descarga. Se rechazan las claves que
contienen segmentos de punto, barras invertidas, rutas absolutas o un enlace simbólico existente
en el destino. Una clave de archivo como reports también entra en conflicto
con una clave anidada como reports/january.txt; asigna nombres distintos a esos objetos
en el bucket.
Si el inicio falla, comprueba si los puertos 9000 o 9001 están ocupados. Si el cliente no puede acceder al bucket, confirma que el servidor siga en ejecución y que la carga inicial haya tenido éxito. Un fallo de permisos al listar o acceder a un objeto devuelve un estado distinto de cero. Un error de permisos en el destino o un conflicto entre archivo y directorio también hace fallar la ejecución; las descargas completadas por otros workers permanecen en el disco.
Prueba tu implementación
Guarda el punto de entrada siguiente como minio-sync/main.go, junto a
client.go y sync.go. Su contexto de señales cancela las
solicitudes en curso al recibir Ctrl+C o SIGTERM. El mensaje de éxito solo se imprime después de
que el listado y todos los workers hayan terminado sin errores:
// main.go
package main
import (
"context"
"log"
"os/signal"
"syscall"
)
func main() {
// Ctrl-C cancels the context, which stops the workers and makes downloadFiles
// report the interruption instead of exiting as if the sync had finished.
ctx, stop := signal.NotifyContext(context.Background(), syscall.SIGINT, syscall.SIGTERM)
defer stop()
// Create MinIO client
client, err := createMinioClient(ctx)
if err != nil {
log.Fatalf("Failed to create MinIO client: %v", err)
}
// Start download. bucketName comes from client.go.
if err := downloadFiles(ctx, client, bucketName, "./downloads"); err != nil {
log.Fatalf("Download failed: %v", err)
}
log.Println("Download completed successfully")
}
Ejecuta el descargador desde el directorio padre:
(
cd minio-sync &&
GOWORK=off GOTOOLCHAIN=local go run .
)
Registra Download completed successfully y escribe los archivos en minio-sync/downloads/.
En este bucket de ejemplo recién creado, verifica cada ruta y cada byte con:
(
cd minio-sync &&
diff -r sample-input downloads
)
La ausencia de salida y un estado cero indican que los directorios coinciden, incluidos los
archivos vacíos y binarios. Un byte distinto o un archivo faltante hace que
diff falle. Los 303 objetos superan la capacidad de ambos búferes de canal
de 100 entradas. Ejecuta de nuevo el descargador para reemplazar las copias locales; vuelve a
descargar todos los objetos. Los archivos que no corresponden a las claves actuales del bucket
permanecen en downloads/.
Consideraciones de seguridad
Reserva downloads/ exclusivamente para este proceso. Las comprobaciones de rutas
rechazan los enlaces simbólicos existentes, pero no impiden que otro proceso cambie una ruta entre
su inspección y la escritura. No constituyen un entorno aislado para el sistema de archivos.
Los directorios nuevos usan el modo 0700 y los archivos temporales usan
0600; los permisos de los directorios existentes no se hacen más
restrictivos. Ejecuta solo un descargador a la vez sobre ese directorio.
Para un servicio remoto, reemplaza las credenciales del entorno de pruebas por una cuenta limitada a las operaciones necesarias del bucket y usa HTTPS. Las credenciales locales fijas de root de este ejemplo son solo para el servidor desechable. Un servicio de almacenamiento con soporte y sus requisitos de despliegue son independientes de este laboratorio con el servidor comunitario archivado.
Reflexiones finales
Cada descarga exitosa reemplaza un archivo local mediante un archivo temporal en el mismo directorio y un cambio de nombre en el mismo sistema de archivos local. Una ejecución fallida puede aun así contener descargas exitosas; no revierte el bucket. El programa no compara marcas de tiempo, no omite objetos sin cambios, no elimina archivos locales obsoletos ni crea una instantánea de un bucket que está cambiando. Mantén el bucket estable durante una ejecución cuando necesites que el resultado represente un conjunto coherente de objetos.
